Skip to content

deps: replace cloudflare/zlib with zlib-ng 2.3.3 - #29433

Merged
Jarred-Sumner merged 7 commits into
mainfrom
claude/zlib-ng
Apr 18, 2026
Merged

Jarred-Sumner merged 7 commits into
mainfrom
claude/zlib-ng

Conversation

@Jarred-Sumner

Copy link
Copy Markdown
Collaborator

What does this PR do?

Replaces the cloudflare/zlib fork (last commit Oct 2023) with zlib-ng 2.3.3 in ZLIB_COMPAT mode. zlib-ng is actively maintained, ships in Node 24+ and Chromium, and provides runtime-dispatched SIMD across AVX-512/AVX2/SSE2/NEON/SVE/RVV for CRC32, adler32, longest-match, and chunk-copy.

Supersedes #16100, #8529.

Benchmarks

Xeon Platinum 8375C (Ice Lake, AVX-512), linux-x64 release build vs system bun 1.3.13. Run with bench/snippets/zlib-comprehensive.mjs and bench/snippets/zlib.mjs (both included).

Operation cloudflare zlib-ng Speedup
gzipSync html-128K L1 275 µs 107 µs 2.59x
gzipSync html-1M L1 2.23 ms 892 µs 2.50x
gzipSync json-128K L6 897 µs 483 µs 1.86x
deflate 123K L6 (async) 373 µs 68 µs 5.48x
gunzipSync html-1M 561 µs 522 µs 1.07x
gunzipSync binary-128K 31.6 µs 26.7 µs 1.18x
createGzip stream L1 1M 3.76 ms 2.68 ms 1.40x
createGunzip stream 1M 1.24 ms 1.18 ms 1.05x
fetch() 11KB gzip decode 42.9 µs 41.6 µs parity
gzipSync 13B (init overhead) 5.04 µs 7.12 µs 0.71x

The streaming-inflate regression that blocked #16100 (Jan 2025, zlib-ng pre-2.2) does not reproduce on 2.3.3. The only downside is ~2µs higher per-stream init cost from larger state structs, amortized away on payloads ≥4KB.

Compression ratio at level=6 is +0.4% vs cloudflare (different match-finding heuristics). Negligible.

Security hardening

Built with -DWITH_INFLATE_STRICT=ON. zlib-ng commit 340f2f6e moved inflateBack()'s distance-too-far-back check behind a default-off #ifdef; upstream zlib has it unconditional. Bun doesn't call inflateBack(), but this hardens against heap OOB reads on malicious raw-deflate with windowBits<15 for anything else linking the same lib, at zero cost to inflate() proper.

Why pin to 2.3.3 (not develop)

Two regressions landed on zlib-ng develop after 2.3.3 that are not present at this commit (documented in zlib.ts):

  • 172b8544 — inverted COPY guard disables Chorba CRC32 fast-path on PCLMULQDQ-only x64
  • e5129cfe — deflateBound() hits __builtin_unreachable() after Z_FINISH

Re-audit before bumping past 2.3.3.

Build system changes

zlib-ng generates zlib.h at cmake-configure time into the build dir (it doesn't exist in source). This required:

  • provides.includes → depBuildDir(cfg, "zlib") instead of source dir
  • libarchive's -I → build dir
  • fetchDeps now resolves to the cross-dep's build outputs (lib files) instead of just the source .ref stamp, so libarchive's configure waits for zlib's configure to have run. resolveDep() takes a map of previously-resolved deps.

Drops 4 cloudflare-specific vendor patches.

How did you verify your code works?

  • linux-x64 release build: bun run build:release clean → smoke test passes
  • test/js/node/zlib/zlib.test.js: 376 pass, 0 fail (release build)
  • bun bd test test/js/node/zlib/: deflate/gzip/inflate tests pass (1 unrelated brotli timeout in debug — createBrotliCompress slowness, untouched by this PR)
  • Build-graph ordering verified: build.ninja shows libarchive configure has deps/zlib/libz.a as order-only input
  • bunx tsc --noEmit -p scripts/build/tsconfig.json clean
  • Windows (lib name → zlibstatic) — needs CI
  • aarch64/musl — needs CI

🤖 Generated with Claude Code

cloudflare/zlib has had no commits since Oct 2023. zlib-ng is actively
maintained, ships in Node 24+ and Chromium, and provides runtime-dispatched
SIMD (AVX-512/AVX2/NEON/SVE/RVV) for CRC32, adler32, hash, and chunk ops.

Benchmark results on Xeon 8375C (Ice Lake, AVX-512), linux-x64 release:
- gzipSync level=1: 1.8x-2.6x faster
- gzipSync level=6: 1.1x-1.9x faster
- deflate 123K level=6: 5.5x faster (373µs -> 68µs)
- gunzipSync: 3-18% faster
- createGzip stream: 1.13x-1.40x faster
- fetch() gzip decode: parity (the #16100 blocker no longer reproduces)
- Init overhead on 13B payload: +2µs (+40%), amortized away on >=4KB

Build with WITH_INFLATE_STRICT=ON to harden inflateBack() against a heap
OOB read on malicious raw-deflate with windowBits<15 (zlib-ng 340f2f6e
moved this check behind a default-off ifdef; Bun doesn't call inflateBack
but defense-in-depth is free here).

Pinned to release tag 2.3.3 — two regressions on develop (172b8544 inverted
CRC32 COPY guard, e5129cfe deflateBound UB) are NOT present at this commit.

Build system: zlib-ng generates zlib.h at configure time into the build dir,
so fetchDeps now resolves to the cross-dep's build outputs (not just source
stamp) and libarchive's -I points at depBuildDir. Drops 4 vendor patches.

Supersedes #16100, #8529.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@robobun

robobun commented Apr 18, 2026 •

Copy link
Copy Markdown
Collaborator
Updated 1:17 AM PT - Apr 18th, 2026

Your commit d1d7fdc is building: #46231

@github-actions

Copy link
Copy Markdown
Contributor

Found 3 issues this PR may fix:

  1. deflateSync produces binary that inflateSync fails to decompress #8886 - deflateSync with level 9 and windowBits -15 produces data that inflateSync fails to decompress ("invalid stored block lengths"); likely a bug in the Cloudflare zlib fork's deflate implementation
  2. Decompression error: ZlibError - empty chunked gzip response breaks fetch() #23149 - Empty chunked gzip response causes ZlibError during fetch() decompression; zlib-ng's stricter inflate implementation may handle this edge case correctly
  3. Safari WebSocket crashes with perMessageDeflate when mixing server.publish() and ws.send() #25673 - WebSocket perMessageDeflate crashes when mixing server.publish() and ws.send(); deflate context corruption may stem from bugs in the Cloudflare zlib fork's state management

If this is helpful, copy the block below into the PR description to auto-close these issues on merge.

Fixes #8886
Fixes #23149
Fixes #25673

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Apr 18, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Switches vendored zlib from the Cloudflare fork to zlib-ng (zlib-compat), updates build/config and dependency resolution, removes several zlib patches/workarounds, adds zlib/gzip benchmarks, adjusts Zig/Windows ABI types, and updates docs, licenses, and tests to the new zlib-ng revision.

Changes

Cohort / File(s) Summary
Zlib migration & build config
scripts/build/deps/zlib.ts
Replace upstream from cloudflare/zlib to zlib-ng/zlib-ng, pin new commit, change build to target zlib-ng with fixed CMake args, adjust provided lib names on Windows and include paths to build output.
Removed/updated zlib patches
scripts/build/patches/zlib/remove-machine-x64.patch, patches/zlib/CMakeLists.txt.patch, patches/zlib/ucm.cmake.patch
Delete multiple previous workaround patches (machine:x64 flag and CMake-version tweaks); removed content that altered zlib CMake handling.
New/updated zlib patches for clang-cl
patches/zlib/clang-cl-arm64.patch
Add patch adjusting ARM NEON/ACLE headers and CMake intrinsic detection so clang-cl on Windows ARM64 selects correct intrinsics and avoids MSVC-only paths.
Zlib headers changes
patches/zlib/deflate.h.patch
Remove MSVC-specific likely/unlikely macros and MSVC fallback __builtin_ctzl implementation (deleted MSVC-specific conditional block).
Zig ABI & bindings
src/deps/zlib.win32.zig, src/zlib.zig
Change uLong alias from u64→c_ulong (affects ABI-facing structs and externs), remove exported zlib version constants from Win unit, update extern signatures and cast call sites.
Build system & dependency resolution
scripts/build/source.ts, scripts/build/bun.ts
Redefine fetchDeps semantics to require deps be built before configure; change resolveDep to accept a resolved map (signature change) and ensure zig-only resolution uses isolated maps to avoid shared state.
Dependent build adjustments
scripts/build/deps/libarchive.ts
Switch include path to zlib’s build dir (generated headers); require zlib to be built prior to libarchive configure while keeping ENABLE_ZLIB=OFF and setting HAVE_ZLIB_H=ON.
Benchmarks added
bench/snippets/fetch-gzip.mjs, bench/snippets/zlib-comprehensive.mjs
Add two benchmark scripts: a fetch+gzip server/runner and a comprehensive zlib benchmark exercising sync/async/streaming modes, multiple corpora/levels, and logging of sizes and process.versions.zlib.
Docs & licensing
CLAUDE.md, LICENSE.md, docs/project/license.mdx
Update documentation and license table entries to reference zlib-ng (zlib-compat) instead of the previous cloudflare fork; license label remains zlib.
Tests
test/js/node/process/process.test.js
Update expected process.versions.zlib commit hash to match the new zlib-ng revision.
🚥 Pre-merge checks | ✅ 2
✅ Passed checks (2 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately summarizes the main change: replacing cloudflare/zlib with zlib-ng 2.3.3, which is the core objective of this PR.
Description check ✅ Passed The PR description comprehensively covers both required template sections: it explains what the PR does (dependency replacement, benchmarks, security hardening, build system changes) and how it was verified (linux-x64 release build, zlib tests passing).

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@bench/snippets/fetch-gzip.mjs`:
- Around line 25-33: Add a one-time preflight before calling group("fetch + gzip
decode", ...) that fetches `${base}/gz` and `${base}/plain`, awaits res.text()
for each, and validates the decoded strings match the expected HTML payload (or
that both decode to the same canonical plain response); if they do not match,
throw or fail the run so the benches inside bench(...) do not record misleading
numbers. Use the existing symbols (group, bench, fetch, res.text, base, paths
/gz and /plain) to locate where to insert the preflight and perform the
validation.

In `@bench/snippets/zlib-comprehensive.mjs`:
- Around line 54-165: The benchmark currently discards outputs (especially in
streaming where drain() ignores bytes), so add a correctness warm-up that for
each corpus and for gzipped entries performs a one-time roundtrip check before
running timed groups: for sync cases call zlib.gzipSync/zlib.gunzipSync and
assert output length or hash equals the original; for async helpers (gzip,
gunzip) await a single call and assert equality; for streams
createGzip/createGunzip run pipeline once with Readable.from(chunks) and a drain
that verifies total bytes or computes a checksum, throwing or console.error on
mismatch to fail fast (refer to drain(), gzip, gunzip, zlib.gzipSync,
zlib.gunzipSync, pipeline, streamInputs/streamGzInputs and the bench group
names).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: b7f787d2-6c75-448e-94ef-7a18f4392e65

📥 Commits

Reviewing files that changed from the base of the PR and between 50be3c3 and f839a51.

📒 Files selected for processing (14)
  • CLAUDE.md
  • LICENSE.md
  • bench/snippets/fetch-gzip.mjs
  • bench/snippets/zlib-comprehensive.mjs
  • docs/project/license.mdx
  • patches/zlib/CMakeLists.txt.patch
  • patches/zlib/deflate.h.patch
  • patches/zlib/ucm.cmake.patch
  • scripts/build/bun.ts
  • scripts/build/deps/libarchive.ts
  • scripts/build/deps/zlib.ts
  • scripts/build/patches/zlib/remove-machine-x64.patch
  • scripts/build/source.ts
  • test/js/node/process/process.test.js
💤 Files with no reviewable changes (4)
  • patches/zlib/CMakeLists.txt.patch
  • patches/zlib/ucm.cmake.patch
  • scripts/build/patches/zlib/remove-machine-x64.patch
  • patches/zlib/deflate.h.patch

Comment on lines +25 to +33
group("fetch + gzip decode", () => {
bench("11KB gzipped → text()", async () => {
const res = await fetch(`${base}/gz`);
await res.text();
});
bench("11KB plain → text()", async () => {
const res = await fetch(`${base}/plain`);
await res.text();
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Validate the decoded payload before timing it.

await res.text() only proves the body was consumable. If fetch stops honoring Content-Encoding: gzip and returns compressed bytes, this benchmark still records a number for the wrong path. Add a one-time preflight that /gz and /plain both decode to the original HTML before group() starts.

Suggested guard
 const base = `http://localhost:${server.port}`;
+
+const expectedText = html.toString();
+if ((await fetch(`${base}/gz`).then(res => res.text())) !== expectedText) {
+  throw new Error("gzip benchmark setup is not decoding the response body");
+}
+if ((await fetch(`${base}/plain`).then(res => res.text())) !== expectedText) {
+  throw new Error("plain benchmark setup returned an unexpected body");
+}
 
 group("fetch + gzip decode", () => {
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@bench/snippets/fetch-gzip.mjs` around lines 25 - 33, Add a one-time preflight
before calling group("fetch + gzip decode", ...) that fetches `${base}/gz` and
`${base}/plain`, awaits res.text() for each, and validates the decoded strings
match the expected HTML payload (or that both decode to the same canonical plain
response); if they do not match, throw or fail the run so the benches inside
bench(...) do not record misleading numbers. Use the existing symbols (group,
bench, fetch, res.text, base, paths /gz and /plain) to locate where to insert
the preflight and perform the validation.

Comment on lines +54 to +165
// Pre-compress for decompression benches
const gzipped = {};
for (const [name, buf] of Object.entries(corpora)) {
gzipped[name] = zlib.gzipSync(buf, { level: 6 });
}

// ─── Sync one-shot ───
group("gzipSync level=1", () => {
for (const [name, buf] of Object.entries(corpora)) {
bench(name, () => zlib.gzipSync(buf, { level: 1 }));
}
});

group("gzipSync level=6", () => {
for (const [name, buf] of Object.entries(corpora)) {
bench(name, () => zlib.gzipSync(buf, { level: 6 }));
}
});

group("gunzipSync", () => {
for (const [name, buf] of Object.entries(gzipped)) {
bench(name, () => zlib.gunzipSync(buf));
}
});

// ─── Async one-shot (threadpool) ───
group("gzip async level=6", () => {
for (const [name, buf] of Object.entries(corpora)) {
bench(name, async () => await gzip(buf, { level: 6 }));
}
});

group("gunzip async", () => {
for (const [name, buf] of Object.entries(gzipped)) {
bench(name, async () => await gunzip(buf));
}
});

// ─── Streaming (HTTP server / npm install path) ───
// Feed input in 16KB chunks like a real socket would.
function chunked(buf, size = 16 * 1024) {
const chunks = [];
for (let i = 0; i < buf.length; i += size) chunks.push(buf.subarray(i, i + size));
return chunks;
}

const streamInputs = {
"html-128K": chunked(corpora["html-128K"]),
"html-1M": chunked(corpora["html-1M"]),
};
const streamGzInputs = {
"html-128K": chunked(gzipped["html-128K"]),
"html-1M": chunked(gzipped["html-1M"]),
};

async function drain(stream) {
for await (const _ of stream);
}

group("createGzip stream level=1", () => {
for (const [name, chunks] of Object.entries(streamInputs)) {
bench(name, async () => {
const gz = zlib.createGzip({ level: 1 });
const src = Readable.from(chunks);
await pipeline(src, gz, drain);
});
}
});

group("createGzip stream level=6", () => {
for (const [name, chunks] of Object.entries(streamInputs)) {
bench(name, async () => {
const gz = zlib.createGzip({ level: 6 });
const src = Readable.from(chunks);
await pipeline(src, gz, drain);
});
}
});

group("createGunzip stream", () => {
for (const [name, chunks] of Object.entries(streamGzInputs)) {
bench(name, async () => {
const gz = zlib.createGunzip();
const src = Readable.from(chunks);
await pipeline(src, gz, drain);
});
}
});

// ─── deflateInit/inflateInit overhead (small payloads, many iterations) ───
// zlib-ng has higher init cost due to larger state structs. This matters for
// per-request gzip on tiny responses.
const tiny = Buffer.from("Hello, World!");
const tinyGz = zlib.gzipSync(tiny);

group("init overhead (13B payload)", () => {
bench("gzipSync", () => zlib.gzipSync(tiny, { level: 6 }));
bench("gunzipSync", () => zlib.gunzipSync(tinyGz));
bench("deflateSync", () => zlib.deflateSync(tiny, { level: 6 }));
bench("inflateSync", () => zlib.inflateSync(zlib.deflateSync(tiny)));
});

// ─── Compression ratio (printed, not benched) ───
console.log("\n# Compression ratio (output bytes, level=6):");
console.log("# corpus input output ratio");
for (const [name, buf] of Object.entries(corpora)) {
const out = zlib.gzipSync(buf, { level: 6 });
console.log(
`# ${name.padEnd(12)} ${String(buf.length).padStart(8)} ${String(out.length).padStart(8)} ${((out.length / buf.length) * 100).toFixed(1)}%`,
);
}
console.log(`# zlib version: ${process.versions.zlib}\n`);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Add a correctness check for the benchmark outputs.

These benches discard every result, so truncated or corrupted output still looks like a performance win. The stream case is especially vulnerable because drain() ignores total bytes entirely. Warm each corpus once and assert the decompressed length or hash matches the source before running the timed groups.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@bench/snippets/zlib-comprehensive.mjs` around lines 54 - 165, The benchmark
currently discards outputs (especially in streaming where drain() ignores
bytes), so add a correctness warm-up that for each corpus and for gzipped
entries performs a one-time roundtrip check before running timed groups: for
sync cases call zlib.gzipSync/zlib.gunzipSync and assert output length or hash
equals the original; for async helpers (gzip, gunzip) await a single call and
assert equality; for streams createGzip/createGunzip run pipeline once with
Readable.from(chunks) and a drain that verifies total bytes or computes a
checksum, throwing or console.error on mismatch to fail fast (refer to drain(),
gzip, gunzip, zlib.gzipSync, zlib.gunzipSync, pipeline,
streamInputs/streamGzInputs and the bench group names).

zlib-ng's arch/arm headers and cmake feature probes gate MSVC-specific
includes on _MSC_VER alone. clang-cl defines _MSC_VER but ships clang's
own <arm_neon.h>/<arm_acle.h>, not MSVC's:

- <arm64_neon.h> defines vld1q_u16_x4/vld1q_u8_x4/vst1q_u16_x4 as macros,
  which macro-expand the polyfill function definitions in neon_intrins.h
  into garbage (27 errors per file).
- <intrin.h> on clang-cl lacks __crc32b/__crc32h/__crc32w/__crc32d; those
  live in clang's <arm_acle.h>.

Patch the guards to `_MSC_VER && !__clang__` in both headers and the two
matching cmake check_c_source_compiles probes, so clang-cl takes the
standard ACLE/NEON path and ARM_CRC32_INTRIN/ARM_NEON_HASLD4 get defined.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@patches/zlib/clang-cl-arm64.patch`:
- Line 13: Update the patch metadata to record the upstream tracking reference:
open an issue or PR in the zlib-ng upstream repo describing the clang-cl ARM64
failure and this local fix, then edit the patch header in clang-cl-arm64.patch
(the line that currently reads "Upstream: not yet reported (zlib-ng CI does not
cover clang-cl ARM64).") to include the upstream issue/PR URL and ID plus a
brief one-line summary of the upstream ticket; keep the existing note and
[request_verification] tag so future bumps can determine whether the fix was
merged, reverted, or superseded.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: ec1e8219-3ea4-415e-86b2-c08c62e980bd

📥 Commits

Reviewing files that changed from the base of the PR and between f839a51 and a0f86c5.

📒 Files selected for processing (2)
  • patches/zlib/clang-cl-arm64.patch
  • scripts/build/deps/zlib.ts

so HAVE_ARMV8_INTRIN and NEON_HAS_LD4 succeed under clang-cl, which sets
ARM_CRC32_INTRIN/ARM_NEON_HASLD4 and skips the polyfills entirely.

Upstream: not yet reported (zlib-ng CI does not cover clang-cl ARM64).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

File the upstream issue before this patch becomes institutional memory.

This fix is now load-bearing for a supported toolchain. Please add an upstream issue/PR reference here once filed so the next zlib-ng bump has a clear breadcrumb for whether the patch was merged, reverted, or superseded.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@patches/zlib/clang-cl-arm64.patch` at line 13, Update the patch metadata to
record the upstream tracking reference: open an issue or PR in the zlib-ng
upstream repo describing the clang-cl ARM64 failure and this local fix, then
edit the patch header in clang-cl-arm64.patch (the line that currently reads
"Upstream: not yet reported (zlib-ng CI does not cover clang-cl ARM64).") to
include the upstream issue/PR URL and ID plus a brief one-line summary of the
upstream ticket; keep the existing note and [request_verification] tag so future
bumps can determine whether the fix was merged, reverted, or superseded.

Comment thread scripts/build/deps/zlib.ts
Comment thread scripts/build/source.ts
root and others added 2 commits April 18, 2026 06:59
cloudflare/zlib typedef'd uLong as uint64_t (8 bytes everywhere). zlib-ng
compat mode (and stock zlib) use `unsigned long` — 4 bytes on Windows LLP64.
Bun's Zig bindings hardcoded uLong = u64, so @sizeof(z_stream) on Windows
was 112 vs zlib-ng's 88, tripping CHECK_VER_STSIZE in deflateInit_/inflateInit_
and returning Z_VERSION_ERROR. Downstream code that didn't check (s3) then
deref'd a null state pointer.

uLong = c_ulong matches the C side on all platforms. Version string is not
the cause: all callers pass zlibVersion() and the check only compares [0].
Dropped the unused ZLIB_VERSION/ZLIB_VERNUM constants from the win32 binding
since they were stale and misleading.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
zlib-ng sets CMAKE_DEBUG_POSTFIX "d" inside if(MSVC), which clang-cl
satisfies, so debug builds produce zlibstaticd.lib.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@scripts/build/source.ts`:
- Around line 377-385: The fetchDeps resolution currently uses r.outputs (final
build artifacts) for nested-CMake dependencies causing downstream configure to
wait on final libs instead of configure-time artifacts; change the dependency
model so each dep exposes two outputs: a configure-stage stamp (e.g.,
r.configureStamp or r.configureOutput) produced by the cmake configure step and
the existing final outputs (r.outputs) produced by build. Update the logic in
fetchDeps resolution (the code that maps deps to r.outputs) to use the
configure-stage stamp as the implicit/Order-only input for downstream configure
runs while preserving r.outputs as the implicit/explicit build-time dependency,
and ensure zlib-ng and other nestedCMake deps create and export that
configure-stage stamp (and that provides.libs remains unchanged for link-time
libraries).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 40804b71-11f6-4d39-9bf2-84250d37878f

📥 Commits

Reviewing files that changed from the base of the PR and between f10b561 and 360b598.

📒 Files selected for processing (2)
  • scripts/build/deps/zlib.ts
  • scripts/build/source.ts

Comment thread scripts/build/source.ts
Comment on lines +377 to +385
* Other deps that must be BUILT before this dep's configure runs.
* Used for header-level dependencies — e.g. libarchive needs zlib's
* headers at compile time (`-I${vendorDir}/zlib`), so zlib must be
* fetched first. This adds an order-only dep on the other dep's source
* stamp — it does NOT link the other dep's libs (that's `provides.libs`).
* headers at configure time (`check_include_file("zlib.h")`). zlib-ng
* generates `zlib.h` during its own cmake configure, so libarchive must
* wait for zlib's full build, not just its source fetch.
*
* Resolves to the named dep's build outputs (lib files for nested-cmake,
* source stamp for header-only). Order-only on configure, implicit on
* build. Does NOT link the other dep's libs (that's `provides.libs`).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Track configure-time artifacts separately from final build outputs.

Line 710 now resolves fetchDeps to r.outputs, but for nested-CMake deps those are the final lib*.a files, not the configure-time contract. For zlib-ng, zlib.h is produced during cmake -B, so downstream configure currently waits for the full zlib build and still does not get dirtied when that generated header changes. That leaves room for stale CMakeCache.txt results across dep bumps or branch switches.

Please expose a configure-stage output/stamp and use that as an implicit input to downstream configure, while keeping final libs as the build-time dependency.

Possible direction
 export interface ResolvedDep {
+  configureOutputs: string[];
   outputs: string[];
 }

- const fetchDepStamps = (dep.fetchDeps ?? []).flatMap(d => {
+ const fetchDepConfigureInputs = (dep.fetchDeps ?? []).flatMap(d => {
    const r = resolved.get(d);
    assert(r, `${dep.name}: fetchDeps references '${d}' but it wasn't resolved first — fix allDeps ordering`);
-   return r.outputs;
+   return r.configureOutputs;
  });

Also applies to: 697-714

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@scripts/build/source.ts` around lines 377 - 385, The fetchDeps resolution
currently uses r.outputs (final build artifacts) for nested-CMake dependencies
causing downstream configure to wait on final libs instead of configure-time
artifacts; change the dependency model so each dep exposes two outputs: a
configure-stage stamp (e.g., r.configureStamp or r.configureOutput) produced by
the cmake configure step and the existing final outputs (r.outputs) produced by
build. Update the logic in fetchDeps resolution (the code that maps deps to
r.outputs) to use the configure-stage stamp as the implicit/Order-only input for
downstream configure runs while preserving r.outputs as the implicit/explicit
build-time dependency, and ensure zlib-ng and other nestedCMake deps create and
export that configure-stage stamp (and that provides.libs remains unchanged for
link-time libraries).

Comment thread src/zlib.zig
root and others added 2 commits April 18, 2026 07:21
Missed in f10b561. On Windows LLP64 the C function returns 32-bit EAX but
Zig was reading 64-bit RAX, passing upper-bit garbage to
ensureTotalCapacityPrecise. Also @intcast the usize -> uLong call site.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Export uLong/uLongf from src/zlib.zig so callers can use the portable type
- HashObject crc32: local accumulator was u64, now bun.zlib.uLong
- crash_handler compress2: len was *usize, now *bun.zlib.uLong; @intcast
  the message.len source and the slice index

zig:check-all clean on all 14 platform/mode combos.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/zlib.zig`:
- Line 847: The call to deflateBound uses `@intCast`(input.len) which can truncate
a 64-bit usize to 32-bit uLong on LLP64 (Windows); add an explicit bounds guard
before narrowing input.len to uLong in the
zlib_reader.list.ensureTotalCapacityPrecise call: check that input.len <=
`@as`(usize, `@intCast`(usize, `@as`(u64, `@as`(uLong, -1)))) or simply compare against
the max value of uLong, and return or error if it exceeds that limit; then
safely cast input.len to uLong when calling deflateBound(&zlib_reader.zlib,
`@intCast`(uLong, input.len)) to prevent silent truncation and potential heap
corruption.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6f749c96-9acb-406d-a2df-ab366d0be92a

📥 Commits

Reviewing files that changed from the base of the PR and between 360b598 and 8bc0ff7.

📒 Files selected for processing (1)
  • src/zlib.zig

Comment thread src/zlib.zig
)) {
ReturnCode.Ok => {
zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, input.len)) catch {
zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, @intCast(input.len))) catch {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify deflateBound usage and whether narrowing casts are guarded.
# Expected: each usize->uLong cast for deflateBound has an explicit pre-check.

rg -nP --type=zig -C3 '\bdeflateBound\s*\('
rg -nP --type=zig -C4 '@intCast\s*\(\s*input\.len\s*\)'
rg -nP --type=zig -C6 'input\.len\s*>\s*[^;\n]+'

Repository: oven-sh/bun

Length of output: 21212


🏁 Script executed:

# Check uLong definition and platform-specific sizing
rg -nP 'typedef.*uLong|const.*uLong|pub\s+const.*uLong' --type=zig --type=c | head -20

# Look for uLong definition in C headers
fd -e h | xargs grep -l 'uLong' 2>/dev/null | head -10

Repository: oven-sh/bun

Length of output: 1872


🏁 Script executed:

# Examine the context around all uLong casts in src/zlib.zig
sed -n '830,860p' src/zlib.zig

Repository: oven-sh/bun

Length of output: 1170


🏁 Script executed:

# Check for other input.len casts in the file - specifically to smaller integer types
rg -nP '@intCast\s*\(\s*input\.len\s*\)|@as\s*\(\s*u(Int|Long)' src/zlib.zig -A2 -B2

Repository: oven-sh/bun

Length of output: 623


🏁 Script executed:

# Look for existing bounds guards on input.len casts in the wider codebase
rg -nP 'input\.len\s*>\s*.*\n.*@intCast|if.*input\.len.*\n.*@intCast' --multiline --type=zig | head -15

Repository: oven-sh/bun

Length of output: 37


🏁 Script executed:

# Check what error types are used for input validation in zlib.zig
rg -nP '(return error\.|error\.)' src/zlib.zig | head -20

# Check if error.InvalidArgument is used elsewhere in codebase for input validation
rg -nP 'error\.InvalidArgument' --type=zig | head -10

Repository: oven-sh/bun

Length of output: 2911


🏁 Script executed:

# Verify the max value computation is correct for uLong on different platforms
cat > /tmp/test_ulong.zig << 'EOF'
const std = `@import`("std");
const c = `@cImport`({
    `@cInclude`("limits.h");
});

pub fn main() void {
    var writer = std.io.getStdOut().writer();
    // Simulating what ~@as(uLong, 0) would compute
    const max_u32: u32 = ~@as(u32, 0);
    const max_u64: u64 = ~@as(u64, 0);
    writer.print("max u32: {}\n", .{max_u32}) catch {};
    writer.print("max u64: {}\n", .{max_u64}) catch {};
}
EOF
# This shows the logic is sound for computing max values

Repository: oven-sh/bun

Length of output: 37


🏁 Script executed:

# Check if there are similar checks elsewhere in zlib.zig or related files
rg -nP 'if.*\.len\s*>\s*' src/zlib.zig src/deps/zlib.*.zig -B2 -A3

Repository: oven-sh/bun

Length of output: 501


🏁 Script executed:

# Specifically verify: are lines 171-172 in same function scope as line 847?
sed -n '160,180p' src/zlib.zig

Repository: oven-sh/bun

Length of output: 836


Add an explicit bounds guard before narrowing input.len to uLong at line 847.

On Windows (LLP64 platform), uLong is 32-bit while usize is 64-bit. Direct casting via @intCast(input.len) silently truncates inputs larger than 4GB, causing deflateBound() to compute an undersized buffer allocation and risking heap corruption.

Proposed fix
         )) {
             ReturnCode.Ok => {
-                zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, `@intCast`(input.len))) catch {
+                const max_source_len: usize = `@intCast`(~@as(uLong, 0));
+                if (input.len > max_source_len) {
+                    zlib_reader.deinit();
+                    return error.InvalidArgument;
+                }
+                const source_len: uLong = `@intCast`(input.len);
+                zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, source_len)) catch {
                     zlib_reader.deinit();
                     return error.OutOfMemory;
                 };
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, @intCast(input.len))) catch {
const max_source_len: usize = `@intCast`(~@as(uLong, 0));
if (input.len > max_source_len) {
zlib_reader.deinit();
return error.InvalidArgument;
}
const source_len: uLong = `@intCast`(input.len);
zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, source_len)) catch {
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/zlib.zig` at line 847, The call to deflateBound uses `@intCast`(input.len)
which can truncate a 64-bit usize to 32-bit uLong on LLP64 (Windows); add an
explicit bounds guard before narrowing input.len to uLong in the
zlib_reader.list.ensureTotalCapacityPrecise call: check that input.len <=
`@as`(usize, `@intCast`(usize, `@as`(u64, `@as`(uLong, -1)))) or simply compare against
the max value of uLong, and return or error if it exceeds that limit; then
safely cast input.len to uLong when calling deflateBound(&zlib_reader.zlib,
`@intCast`(uLong, input.len)) to prevent silent truncation and potential heap
corruption.

Comment thread src/zlib.zig
@Jarred-Sumner

Copy link
Copy Markdown
Collaborator Author

@robobun fix

❌ CPU instruction violation on Windows x64 — 1 check(s) failed
The baseline build contains instructions not available on Nehalem (SSE4.2, no AVX/AVX2/AVX512).

Static instruction scan
Static scan violations
adler32_avx2 [AVX, AVX2] (87 insns)
0x0142fc3444 Vmovdqa (AVX)
0x0142fc344a Vmovdqa (AVX)
0x0142fc3450 Vmovdqa (AVX)
... 84 more
adler32_avx512 [AVX, AVX512BW, AVX512F, AVX512VL] (68 insns)
0x0142fc56c3 Vpxor (AVX)
0x0142fc56c7 Vmovdqa64 (AVX512F)
0x0142fc56d1 Vmovdqa64 (AVX512F)
... 65 more
adler32_avx512_vnni [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, AVX512_VNNI] (68 insns)
0x0142fc69fd Vpxor (AVX)
0x0142fc6a01 Vmovdqa64 (AVX512F)
0x0142fc6a10 Vmovdqa64 (AVX512F)
... 65 more
adler32_fold_copy_avx2 [AVX, AVX2] (92 insns)
0x0142fc37d5 Vmovdqa (AVX)
0x0142fc37db Vmovdqa (AVX)
0x0142fc37e1 Vmovdqa (AVX)
... 89 more
adler32_fold_copy_avx512 [AVX, AVX512BW, AVX512F, AVX512VL, BMI2] (75 insns)
0x0142fc5454 Vpxor (AVX)
0x0142fc5458 Vmovdqa64 (AVX512F)
0x0142fc5462 Vmovdqa64 (AVX512F)
... 72 more
adler32_fold_copy_avx512_vnni [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, AVX512_VNNI, BMI2] (119 insns)
0x0142fc6cb6 Vpxor (AVX)
0x0142fc6cba Vmovdqa (AVX)
0x0142fc6cc7 Vmovdqa (AVX)
... 116 more
chunkmemset_avx2 [AVX, AVX2, BMI2] (38 insns)
0x0142fc1a04 Vzeroupper (AVX)
0x0142fc1a15 Vmovdqu (AVX)
0x0142fc1a19 Vmovdqu (AVX)
... 35 more
chunkmemset_avx512 [AVX, AVX2, AVX512BW, AVX512VL, BMI2] (66 insns)
0x0142fc456c Vmovdqu (AVX)
0x0142fc4570 Vmovdqu (AVX)
0x0142fc45b5 Vzeroupper (AVX)
... 63 more
compare256_avx2 [AVX, AVX2] (27 insns)
0x0142fc0e30 Vmovdqu (AVX)
0x0142fc0e34 Vpcmpeqb (AVX2)
0x0142fc0e38 Vpmovmskb (AVX2)
... 24 more
compare256_avx512 [AVX, AVX512BW, AVX512F, AVX512VL] (26 insns)
0x0142fc3bb0 Vmovdqu (AVX)
0x0142fc3bb4 Vpcmpeqb (AVX512VL)
0x0142fc3bb4 Vpcmpeqb (AVX512BW)
... 23 more
crc32_fold_pclmulqdq [PCLMULQDQ] (100 insns)
0x0142fbe5b5 Pclmulqdq (PCLMULQDQ)
0x0142fbe5bc Pclmulqdq (PCLMULQDQ)
0x0142fbe677 Pclmulqdq (PCLMULQDQ)
... 97 more
crc32_fold_pclmulqdq_copy [PCLMULQDQ] (24 insns)
0x0142fbf3b0 Pclmulqdq (PCLMULQDQ)
0x0142fbf3bb Pclmulqdq (PCLMULQDQ)
0x0142fbf414 Pclmulqdq (PCLMULQDQ)
... 21 more
crc32_fold_pclmulqdq_final [PCLMULQDQ] (10 insns)
0x0142fbf780 Pclmulqdq (PCLMULQDQ)
0x0142fbf790 Pclmulqdq (PCLMULQDQ)
0x0142fbf7a3 Pclmulqdq (PCLMULQDQ)
... 7 more
crc32_fold_vpclmulqdq [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, PCLMULQDQ, VPCLMULQDQ] (317 insns)
0x0142fc58fc Vmovaps (AVX)
0x0142fc5905 Vmovaps (AVX)
0x0142fc590e Vmovdqa (AVX)
... 314 more
crc32_fold_vpclmulqdq_copy [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, PCLMULQDQ, VPCLMULQDQ] (271 insns)
0x0142fc607e Vmovaps (AVX)
0x0142fc6087 Vmovaps (AVX)
0x0142fc60ae Vpxor (AVX)
... 268 more
crc32_fold_vpclmulqdq_final [AVX, AVX2, AVX512F, AVX512VL, PCLMULQDQ] (42 insns)
0x0142fc6750 Vpbroadcastq (AVX2)
0x0142fc6759 Vmovdqa (AVX)
0x0142fc675d Vpclmulqdq (PCLMULQDQ)
... 39 more
... 50 more lines (see step log)
If these are runtime-dispatched behind a CPUID gate: add each symbol to scripts/verify-baseline-static/allowlist-x64-windows.txt with a comment pointing at the gate. Feature ceilings (the [FEAT, ...] bracket) should list what the gate checks.

If there's no gate: this is a real bug — a -march leaked into a subbuild. Find the translation unit and fix its compile flags.

❌ CPU instruction violation on Linux x64 — 1 check(s) failed
The baseline build contains instructions not available on Nehalem (SSE4.2, no AVX/AVX2/AVX512).

Static instruction scan
Static scan violations
adler32_avx2 [AVX, AVX2] (70 insns)
0x0004066d60 Vpxor (AVX)
0x0004066d64 Vmovdqa (AVX)
0x0004066d6c Vmovdqa (AVX)
... 67 more
adler32_avx512 [AVX, AVX2, AVX512BW, AVX512F, AVX512VL] (70 insns)
0x00040676dc Vpxor (AVX)
0x00040676e0 Vmovdqa64 (AVX512F)
0x00040676ea Vpbroadcastd (AVX512F)
... 67 more
adler32_avx512_vnni [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, AVX512_VNNI] (88 insns)
0x0004069080 Vpxor (AVX)
0x0004069084 Vmovdqa64 (AVX512F)
0x0004069093 Vmovdqa64 (AVX512F)
... 85 more
adler32_fold_copy_avx2 [AVX, AVX2] (75 insns)
0x0004067081 Vpxor (AVX)
0x0004067085 Vmovdqa (AVX)
0x000406708d Vmovdqa (AVX)
... 72 more
adler32_fold_copy_avx512 [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, BMI2] (79 insns)
0x00040673ec Vpxor (AVX)
0x00040673f0 Vmovdqa64 (AVX512F)
0x00040673fa Vpbroadcastd (AVX512F)
... 76 more
adler32_fold_copy_avx512_vnni [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, AVX512_VNNI, BMI2] (84 insns)
0x00040693ea Vpxor (AVX)
0x00040693ee Vmovdqa (AVX)
0x00040693fb Vmovdqa (AVX)
... 81 more
chunkcopy_safe [AVX] (22 insns)
0x0004065fa2 Vzeroupper (AVX)
0x0004065fd0 Vmovups (AVX)
0x0004065fd4 Vmovups (AVX)
... 19 more
chunkmemset_avx2 [AVX, AVX2, BMI2] (38 insns)
0x000406500e Vzeroupper (AVX)
0x000406501e Vmovdqu (AVX)
0x0004065022 Vmovdqu (AVX)
... 35 more
chunkmemset_avx512 [AVX, AVX2, AVX512BW, AVX512VL, BMI2] (66 insns)
0x00040679bb Vmovdqu (AVX)
0x00040679bf Vmovdqu (AVX)
0x0004067a02 Bzhi (BMI2)
... 63 more
compare256_avx2 [AVX, AVX2] (27 insns)
0x0004066214 Vmovdqu (AVX)
0x0004066218 Vpcmpeqb (AVX2)
0x000406621c Vpmovmskb (AVX2)
... 24 more
compare256_avx512 [AVX, AVX512BW, AVX512F, AVX512VL] (26 insns)
0x0004068644 Vmovdqu (AVX)
0x0004068648 Vpcmpeqb (AVX512VL)
0x0004068648 Vpcmpeqb (AVX512BW)
... 23 more
crc32_fold_pclmulqdq [PCLMULQDQ] (100 insns)
0x000406372d Pclmulqdq (PCLMULQDQ)
0x0004063734 Pclmulqdq (PCLMULQDQ)
0x00040637f8 Pclmulqdq (PCLMULQDQ)
... 97 more
crc32_fold_pclmulqdq_copy [PCLMULQDQ] (24 insns)
0x00040646d6 Pclmulqdq (PCLMULQDQ)
0x00040646e1 Pclmulqdq (PCLMULQDQ)
0x0004064737 Pclmulqdq (PCLMULQDQ)
... 21 more
crc32_fold_pclmulqdq_final [PCLMULQDQ] (10 insns)
0x0004064b04 Pclmulqdq (PCLMULQDQ)
0x0004064b14 Pclmulqdq (PCLMULQDQ)
0x0004064b27 Pclmulqdq (PCLMULQDQ)
... 7 more
crc32_fold_vpclmulqdq [AVX, AVX2, AVX512F, AVX512VL, PCLMULQDQ, VPCLMULQDQ] (391 insns)
0x000406969d Vpxor (AVX)
0x00040696a1 Vmovdqa (AVX)
0x00040696aa Vmovdqa (AVX)
... 388 more
crc32_fold_vpclmulqdq_copy [AVX, AVX2, AVX512F, AVX512VL, PCLMULQDQ, VPCLMULQDQ] (309 insns)
0x000406a0f5 Vpxor (AVX)
0x000406a0f9 Vmovdqa (AVX)
0x000406a102 Vmovdqa (AVX)
... 306 more
... 55 more lines (see step log)
If these are runtime-dispatched behind a CPUID gate: add each symbol to scripts/verify-baseline-static/allowlist-x64.txt with a comment pointing at the gate. Feature ceilings (the [FEAT, ...] bracket) should list what the gate checks.

If there's no gate: this is a real bug — a -march leaked into a subbuild. Find the translation unit and fix its compile flags.

❌ CPU instruction violation on Linux x64 — 1 check(s) failed
The baseline build contains instructions not available on Nehalem (SSE4.2, no AVX/AVX2/AVX512).

Static instruction scan
Static scan violations
adler32_avx2 [AVX, AVX2] (70 insns)
0x00040cad30 Vpxor (AVX)
0x00040cad34 Vmovdqa (AVX)
0x00040cad3c Vmovdqa (AVX)
... 67 more
adler32_avx512 [AVX, AVX2, AVX512BW, AVX512F, AVX512VL] (70 insns)
0x00040cb6ac Vpxor (AVX)
0x00040cb6b0 Vmovdqa64 (AVX512F)
0x00040cb6ba Vpbroadcastd (AVX512F)
... 67 more
adler32_avx512_vnni [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, AVX512_VNNI] (88 insns)
0x00040cd2b0 Vpxor (AVX)
0x00040cd2b4 Vmovdqa64 (AVX512F)
0x00040cd2c3 Vmovdqa64 (AVX512F)
... 85 more
adler32_fold_copy_avx2 [AVX, AVX2] (75 insns)
0x00040cb051 Vpxor (AVX)
0x00040cb055 Vmovdqa (AVX)
0x00040cb05d Vmovdqa (AVX)
... 72 more
adler32_fold_copy_avx512 [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, BMI2] (79 insns)
0x00040cb3bc Vpxor (AVX)
0x00040cb3c0 Vmovdqa64 (AVX512F)
0x00040cb3ca Vpbroadcastd (AVX512F)
... 76 more
adler32_fold_copy_avx512_vnni [AVX, AVX2, AVX512BW, AVX512F, AVX512VL, AVX512_VNNI, BMI2] (84 insns)
0x00040cd61a Vpxor (AVX)
0x00040cd61e Vmovdqa (AVX)
0x00040cd62b Vmovdqa (AVX)
... 81 more
chunkmemset_avx2 [AVX, AVX2, BMI2] (37 insns)
0x00040c8851 Vzeroupper (AVX)
0x00040c8861 Vmovdqu (AVX)
0x00040c8865 Vmovdqu (AVX)
... 34 more
chunkmemset_avx512 [AVX, AVX2, AVX512BW, AVX512VL, BMI2] (65 insns)
0x00040cb97e Vmovdqu (AVX)
0x00040cb982 Vmovdqu (AVX)
0x00040cb9b1 Vzeroupper (AVX)
... 62 more
compare256_avx2 [AVX, AVX2] (27 insns)
0x00040ca1e4 Vmovdqu (AVX)
0x00040ca1e8 Vpcmpeqb (AVX2)
0x00040ca1ec Vpmovmskb (AVX2)
... 24 more
compare256_avx512 [AVX, AVX512BW, AVX512F, AVX512VL] (26 insns)
0x00040cc874 Vmovdqu (AVX)
0x00040cc878 Vpcmpeqb (AVX512VL)
0x00040cc878 Vpcmpeqb (AVX512BW)
... 23 more
crc32_fold_pclmulqdq [PCLMULQDQ] (100 insns)
0x00040c718b Pclmulqdq (PCLMULQDQ)
0x00040c7192 Pclmulqdq (PCLMULQDQ)
0x00040c7250 Pclmulqdq (PCLMULQDQ)
... 97 more
crc32_fold_pclmulqdq_copy [PCLMULQDQ] (24 insns)
0x00040c7fc7 Pclmulqdq (PCLMULQDQ)
0x00040c7fd2 Pclmulqdq (PCLMULQDQ)
0x00040c8027 Pclmulqdq (PCLMULQDQ)
... 21 more
crc32_fold_pclmulqdq_final [PCLMULQDQ] (10 insns)
0x00040c8384 Pclmulqdq (PCLMULQDQ)
0x00040c8394 Pclmulqdq (PCLMULQDQ)
0x00040c83a7 Pclmulqdq (PCLMULQDQ)
... 7 more
crc32_fold_vpclmulqdq [AVX, AVX2, AVX512F, AVX512VL, PCLMULQDQ, VPCLMULQDQ] (385 insns)
0x00040cd8bc Vpxor (AVX)
0x00040cd8c0 Vmovdqa (AVX)
0x00040cd8c5 Vmovdqa (AVX)
... 382 more
crc32_fold_vpclmulqdq_copy [AVX, AVX2, AVX512F, AVX512VL, PCLMULQDQ, VPCLMULQDQ] (308 insns)
0x00040ce1c4 Vpxor (AVX)
0x00040ce1c8 Vmovdqa (AVX)
0x00040ce1ce Vmovdqa (AVX)
... 305 more
crc32_fold_vpclmulqdq_final [AVX, AVX2, AVX512F, AVX512VL, PCLMULQDQ] (42 insns)
0x00040ce984 Vmovdqa (AVX)
0x00040ce988 Vpbroadcastq (AVX2)
0x00040ce991 Vpclmulqdq (PCLMULQDQ)
... 39 more
... 50 more lines (see step log)
If these are runtime-dispatched behind a CPUID gate: add each symbol to scripts/verify-baseline-static/allowlist-x64.txt with a comment pointing at the gate. Feature ceilings (the [FEAT, ...] bracket) should list what the gate checks.

If there's no gate: this is a real bug — a -march leaked into a subbuild. Find the translation unit and fix its compile flags.

@robobun

robobun commented Apr 18, 2026 •

Copy link
Copy Markdown
Collaborator

✅ Resolved — zlib-ng functable dispatch targets were added to scripts/verify-baseline-static/allowlist-x64{,-windows}.txt before merge. #29433 is in main.

zlib-ng compiles per-ISA variants (PCLMULQDQ/AVX2/AVX512/VNNI/VPCLMULQDQ)
and selects at runtime via init_functable() -> cpu_check_features() CPUID
in vendor/zlib/functable.c. These are intentionally present in baseline
builds; the static scanner just needs to know they're gated.

Ceilings match the gate predicates in arch/x86/x86_features.c:
  has_pclmulqdq, has_avx2 && has_bmi2, has_avx512_common (F+DQ+BW+VL+BMI2),
  has_avx512vnni, has_vpclmulqdq.

chunkcopy_safe is a static-inline (inflate_p.h) that the compiler may
outline from one of the per-ISA inffast_tpl.h instantiations; the outlined
local copy is only reachable from the gated inflate_fast_* of that TU.

Replaces the old single-symbol cloudflare-zlib `crc32` entry on Linux.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@Jarred-Sumner
Jarred-Sumner merged commit 8f2519f into main Apr 18, 2026
3 of 5 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the claude/zlib-ng branch April 18, 2026 08:18
Comment thread src/zlib.zig
)) {
ReturnCode.Ok => {
zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, input.len)) catch {
zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, @intCast(input.len))) catch {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 At src/zlib.zig:847, deflateBound is called as deflateBound(&zlib_reader.zlib, @intCast(input.len)), but this PR changed uLong from u64 to c_ulong (= u32 on Windows LLP64), so @intCast will panic in Zig debug/safe builds for any input.len > 4GB. The surrounding code at lines 817–818 uses @truncate for the exact same narrowing cast (avail_in, total_in); line 847 should match. Fix: change @intCast(input.len) → @truncate(input.len).

Extended reasoning...

What the bug is and how it manifests

This PR correctly changes const uLong = c_ulong at src/zlib.zig lines 22–26 to fix the Windows LLP64 ABI (where unsigned long is 4 bytes, not 8). However, the deflateBound call site at line 847 was updated to add a cast but used @intCast instead of @truncate:

zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, @intCast(input.len)))

In Zig, @intCast is a checked narrowing cast: it panics at runtime in debug and ReleaseSafe builds when the source value exceeds the destination type's range. On Windows where uLong = c_ulong = u32, any input.len > 0xFFFF_FFFF (4GB) will trigger a guaranteed runtime panic.

The specific code path that triggers it

ZlibCompressorArrayList.initWithListAllocator (called from ZlibCompressorArrayList.init, which backs synchronous zlib compression) calls deflateBound to pre-size the output buffer. input.len is usize = u64 on all 64-bit platforms; the cast to uLong = u32 on Windows silently works for inputs ≤ 4GB but panics or truncates for larger ones.

Why existing code doesn't prevent it

The two narrowing casts directly above (lines 817–818) for avail_in and total_in — the same usize → uLong narrowing — already use @truncate:

.avail_in = @truncate(input.len),
.total_in = @truncate(input.len),

@truncate is the Zig idiom that explicitly acknowledges truncation without a safety check. Line 847 was missed in this PR when the @intCast was added to satisfy the new type, resulting in an inconsistency that compiles cleanly but behaves differently at runtime.

What the impact would be

  • Debug/ReleaseSafe builds on Windows: guaranteed panic for inputs > 4GB (e.g. any internal Zig caller passing a large []u8 slice).
  • ReleaseFast/ReleaseSmall builds on Windows: silent truncation — deflateBound receives a wrong (too-small) length estimate. The output buffer is under-sized, causing the subsequent deflate() call to return Z_BUF_ERROR (no progress made), breaking synchronous compression.
  • All non-Windows platforms: no impact — c_ulong = u64 there, no truncation occurs.

In practice, JS ArrayBuffer inputs via the public API are capped below 4GB by V8/JSC, so this is unreachable from JavaScript. However, internal Zig callers (or future code) are not subject to that constraint.

How to fix it

Change line 847 to match the pattern used for avail_in/total_in:

// Before:
zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, @intCast(input.len)))
// After:
zlib_reader.list.ensureTotalCapacityPrecise(list_allocator, deflateBound(&zlib_reader.zlib, @truncate(input.len)))

Step-by-step proof

  1. PR changes uLong = c_ulong; on Windows LLP64, c_ulong = u32.
  2. deflateBound C signature: uLong deflateBound(z_streamp strm, uLong sourceLen) — both 32-bit on Windows.
  3. input.len is usize = u64 on 64-bit platforms.
  4. @intCast(input.len) at line 847: if input.len > std.math.maxInt(u32), Zig's safety check fires → panic in debug.
  5. Lines 817–818 use @truncate(input.len) for the identical narrowing — no safety check, explicit truncation.
  6. The inconsistency is the direct result of this PR adding a cast at line 847 but choosing the wrong cast builtin.

🔬 also observed by coderabbitai

@belgattitude

Copy link
Copy Markdown

Thanks for the update. Just upgraded bun yesterday and my ci pipeline failed for regression in size vs nodejs. Is there a chance that the compression ratio has changed between bun 1.3.12 and 1.3.13 ?

  /**
   * Compress the given data and return it as a Uint8Array binary format
   *
   * ```typescript
   * import { Compressor } from '@httpx/compress';
   *
   * const compressor = new Compressor('gzip');
   * const longString = 'Hello, World! 🦆'.repeat(500_000);
   * const compressedBinary = await compressor.Uint8Array(longString);
   * ```
   * @throws Error
   */
  toUint8Array = async <T extends string | Uint8Array>(
    data: T
  ): Promise<Uint8Array> => {
    const readableStream = new ReadableStream({
      start(controller) {
        controller.enqueue(
          typeof data === 'string'
            ? globalCache.utf8TextEncoder.encode(data)
            : data
        );
        controller.close();
      },
    });
    const compressedStream = readableStream.pipeThrough(
      new CompressionStream(this.#algorithm)
    );

    return new Uint8Array(await new Response(compressedStream).arrayBuffer());
  };
image

structwafel pushed a commit to structwafel/bun that referenced this pull request Apr 25, 2026
## What does this PR do?

Replaces the cloudflare/zlib fork (last commit Oct 2023) with
[zlib-ng](https://github.com/zlib-ng/zlib-ng) 2.3.3 in `ZLIB_COMPAT`
mode. zlib-ng is actively maintained, ships in Node 24+ and Chromium,
and provides runtime-dispatched SIMD across
AVX-512/AVX2/SSE2/NEON/SVE/RVV for CRC32, adler32, longest-match, and
chunk-copy.

Supersedes oven-sh#16100, oven-sh#8529.

## Benchmarks

Xeon Platinum 8375C (Ice Lake, AVX-512), linux-x64 release build vs
system bun 1.3.13. Run with `bench/snippets/zlib-comprehensive.mjs` and
`bench/snippets/zlib.mjs` (both included).

| Operation | cloudflare | zlib-ng | Speedup |
|---|---:|---:|---:|
| `gzipSync` html-128K L1 | 275 µs | 107 µs | **2.59x** |
| `gzipSync` html-1M L1 | 2.23 ms | 892 µs | **2.50x** |
| `gzipSync` json-128K L6 | 897 µs | 483 µs | **1.86x** |
| `deflate` 123K L6 (async) | 373 µs | 68 µs | **5.48x** |
| `gunzipSync` html-1M | 561 µs | 522 µs | 1.07x |
| `gunzipSync` binary-128K | 31.6 µs | 26.7 µs | 1.18x |
| `createGzip` stream L1 1M | 3.76 ms | 2.68 ms | **1.40x** |
| `createGunzip` stream 1M | 1.24 ms | 1.18 ms | 1.05x |
| `fetch()` 11KB gzip decode | 42.9 µs | 41.6 µs | parity |
| `gzipSync` 13B (init overhead) | 5.04 µs | 7.12 µs | 0.71x |

The streaming-inflate regression that blocked oven-sh#16100 (Jan 2025, zlib-ng
pre-2.2) **does not reproduce** on 2.3.3. The only downside is ~2µs
higher per-stream init cost from larger state structs, amortized away on
payloads ≥4KB.

Compression ratio at level=6 is +0.4% vs cloudflare (different
match-finding heuristics). Negligible.

## Security hardening

Built with `-DWITH_INFLATE_STRICT=ON`. zlib-ng commit `340f2f6e` moved
`inflateBack()`'s distance-too-far-back check behind a default-off
`#ifdef`; upstream zlib has it unconditional. Bun doesn't call
`inflateBack()`, but this hardens against heap OOB reads on malicious
raw-deflate with `windowBits<15` for anything else linking the same lib,
at zero cost to `inflate()` proper.

## Why pin to 2.3.3 (not develop)

Two regressions landed on zlib-ng `develop` after 2.3.3 that are **not**
present at this commit (documented in `zlib.ts`):
- `172b8544` — inverted `COPY` guard disables Chorba CRC32 fast-path on
PCLMULQDQ-only x64
- `e5129cfe` — `deflateBound()` hits `__builtin_unreachable()` after
`Z_FINISH`

Re-audit before bumping past 2.3.3.

## Build system changes

zlib-ng generates `zlib.h` at cmake-configure time into the **build**
dir (it doesn't exist in source). This required:
- `provides.includes` → `depBuildDir(cfg, "zlib")` instead of source dir
- libarchive's `-I` → build dir
- `fetchDeps` now resolves to the cross-dep's **build outputs** (lib
files) instead of just the source `.ref` stamp, so libarchive's
configure waits for zlib's configure to have run. `resolveDep()` takes a
map of previously-resolved deps.

Drops 4 cloudflare-specific vendor patches.

## How did you verify your code works?

- [x] linux-x64 release build: `bun run build:release` clean → smoke
test passes
- [x] `test/js/node/zlib/zlib.test.js`: **376 pass**, 0 fail (release
build)
- [x] `bun bd test test/js/node/zlib/`: deflate/gzip/inflate tests pass
(1 unrelated brotli timeout in debug — `createBrotliCompress` slowness,
untouched by this PR)
- [x] Build-graph ordering verified: `build.ninja` shows libarchive
configure has `deps/zlib/libz.a` as order-only input
- [x] `bunx tsc --noEmit -p scripts/build/tsconfig.json` clean
- [ ] Windows (lib name → `zlibstatic`) — needs CI
- [ ] aarch64/musl — needs CI

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: root <root@ip-10-0-2-234.us-west-2.compute.internal>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
xhjkl pushed a commit to xhjkl/bun that referenced this pull request May 14, 2026
## What does this PR do?

Replaces the cloudflare/zlib fork (last commit Oct 2023) with
[zlib-ng](https://github.com/zlib-ng/zlib-ng) 2.3.3 in `ZLIB_COMPAT`
mode. zlib-ng is actively maintained, ships in Node 24+ and Chromium,
and provides runtime-dispatched SIMD across
AVX-512/AVX2/SSE2/NEON/SVE/RVV for CRC32, adler32, longest-match, and
chunk-copy.

Supersedes oven-sh#16100, oven-sh#8529.

## Benchmarks

Xeon Platinum 8375C (Ice Lake, AVX-512), linux-x64 release build vs
system bun 1.3.13. Run with `bench/snippets/zlib-comprehensive.mjs` and
`bench/snippets/zlib.mjs` (both included).

| Operation | cloudflare | zlib-ng | Speedup |
|---|---:|---:|---:|
| `gzipSync` html-128K L1 | 275 µs | 107 µs | **2.59x** |
| `gzipSync` html-1M L1 | 2.23 ms | 892 µs | **2.50x** |
| `gzipSync` json-128K L6 | 897 µs | 483 µs | **1.86x** |
| `deflate` 123K L6 (async) | 373 µs | 68 µs | **5.48x** |
| `gunzipSync` html-1M | 561 µs | 522 µs | 1.07x |
| `gunzipSync` binary-128K | 31.6 µs | 26.7 µs | 1.18x |
| `createGzip` stream L1 1M | 3.76 ms | 2.68 ms | **1.40x** |
| `createGunzip` stream 1M | 1.24 ms | 1.18 ms | 1.05x |
| `fetch()` 11KB gzip decode | 42.9 µs | 41.6 µs | parity |
| `gzipSync` 13B (init overhead) | 5.04 µs | 7.12 µs | 0.71x |

The streaming-inflate regression that blocked oven-sh#16100 (Jan 2025, zlib-ng
pre-2.2) **does not reproduce** on 2.3.3. The only downside is ~2µs
higher per-stream init cost from larger state structs, amortized away on
payloads ≥4KB.

Compression ratio at level=6 is +0.4% vs cloudflare (different
match-finding heuristics). Negligible.

## Security hardening

Built with `-DWITH_INFLATE_STRICT=ON`. zlib-ng commit `340f2f6e` moved
`inflateBack()`'s distance-too-far-back check behind a default-off
`#ifdef`; upstream zlib has it unconditional. Bun doesn't call
`inflateBack()`, but this hardens against heap OOB reads on malicious
raw-deflate with `windowBits<15` for anything else linking the same lib,
at zero cost to `inflate()` proper.

## Why pin to 2.3.3 (not develop)

Two regressions landed on zlib-ng `develop` after 2.3.3 that are **not**
present at this commit (documented in `zlib.ts`):
- `172b8544` — inverted `COPY` guard disables Chorba CRC32 fast-path on
PCLMULQDQ-only x64
- `e5129cfe` — `deflateBound()` hits `__builtin_unreachable()` after
`Z_FINISH`

Re-audit before bumping past 2.3.3.

## Build system changes

zlib-ng generates `zlib.h` at cmake-configure time into the **build**
dir (it doesn't exist in source). This required:
- `provides.includes` → `depBuildDir(cfg, "zlib")` instead of source dir
- libarchive's `-I` → build dir
- `fetchDeps` now resolves to the cross-dep's **build outputs** (lib
files) instead of just the source `.ref` stamp, so libarchive's
configure waits for zlib's configure to have run. `resolveDep()` takes a
map of previously-resolved deps.

Drops 4 cloudflare-specific vendor patches.

## How did you verify your code works?

- [x] linux-x64 release build: `bun run build:release` clean → smoke
test passes
- [x] `test/js/node/zlib/zlib.test.js`: **376 pass**, 0 fail (release
build)
- [x] `bun bd test test/js/node/zlib/`: deflate/gzip/inflate tests pass
(1 unrelated brotli timeout in debug — `createBrotliCompress` slowness,
untouched by this PR)
- [x] Build-graph ordering verified: `build.ninja` shows libarchive
configure has `deps/zlib/libz.a` as order-only input
- [x] `bunx tsc --noEmit -p scripts/build/tsconfig.json` clean
- [ ] Windows (lib name → `zlibstatic`) — needs CI
- [ ] aarch64/musl — needs CI

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: root <root@ip-10-0-2-234.us-west-2.compute.internal>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants