Skip to content

node:zlib: reset a zstd stream in place, keep its dictionary and parameters - #44261

Open
robobun wants to merge 4 commits into
mainfrom
robobun/55a00183/zstd-reset-keeps-dictionary-and-params
Open

robobun wants to merge 4 commits into
mainfrom
robobun/55a00183/zstd-reset-keeps-dictionary-and-params

Conversation

@robobun

@robobun robobun commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • reset() of a node:zlib zstd stream drops its dictionary and parameters. After createZstdDecompress({ dictionary }).reset(), a valid frame fails with ZSTD_error_corruption_detected.
  • Context::reset (src/runtime/node/zlib/NativeZstd.rs:461) called init(), which makes a default zstd context. Node did the same until v26.9.0 and keeps the options since v26.10.0 (zlib: fix zstd reset nodejs/node#65867). No user reported this.

Fix

  • Context::reset resets the session on the same context and sets the pledged size again, as Node v26.10.0 does.
  • A one-line zstd patch clears jobReady where zstd erases its job table. Without it, the frame after a reset() in an open frame can crash in ZSTDMT_compressionJob. Node v26.10.0 does.
  • Verified: test/js/node/zlib/zlib-zstd-reset.test.ts and Node's test-zlib-zstd-reset.js. main fails 14 of 17 tests.
  • Self-reviewed: 15 concerns raised, 12 addressed. Notes name the 3 others.

Background

  • A zstd context holds the parameters, a copy of the dictionary and the state of the open frame. A session reset drops only the frame state.
  • With ZSTD_c_nbWorkers, zstd cuts the input into jobs for worker threads. jobReady marks a job that waits for a free worker.
  • Considered: zlib.ts keeps the init arguments and calls init() again. That misses _handle.reset() and adds state to every stream.

Downsides

Notes

Which Node versions keep the options

Node ZstdCompressContext::ResetStream / ZstdDecompressContext::ResetStream
v22.15.0 to v26.9.0, v24.21.0, v22.23.3 return Init(...): a new context, no dictionary, no parameters
v26.10.0 ZSTD_CCtx_reset(session_only) + ZSTD_CCtx_setPledgedSrcSize, ZSTD_DCtx_reset(session_only)

Reproduction

const zlib = require("node:zlib");
const dictionary = Buffer.alloc(4096, "the quick brown fox jumps over the lazy dog ");
const data = Buffer.alloc(2000, "the quick brown fox jumps over the lazy dog ");
const frame = zlib.zstdCompressSync(data, { dictionary });
const d = new zlib.ZstdDecompress({ dictionary });
const h = d._handle; const close = h.close; h.close = () => {};   // keep the handle open between calls
const step = name => {
  try { console.log(name, d._processChunk(frame, zlib.constants.ZSTD_e_end).length); }
  catch (e) { console.log(name, "throws", e.code, e.message); }
  d._handle = h;
};
step("before reset");
h.reset();
step("after reset");
close.call(h);
output
main (542f52b, release) before reset 2000, after reset throws ZSTD_error_corruption_detected Data corruption detected
this PR (release) before reset 2000, after reset 2000
Node v26.10.0 before reset 2000, after reset 2000

The three cases of Node's test (compressor with dictionary, level 19, checksum and pledged size, decompressor with dictionary, decompressor with ZSTD_d_windowLogMax): main gives 76 bytes in place of 26, ZSTD_error_corruption_detected, and 4096 decoded bytes where the limit must reject the frame. This PR and Node v26.10.0 give 26 bytes, the decoded input, and ZSTD_error_frameParameter_windowTooLarge.

The zstd patch

  • ZSTDMT_createCompressionJob() sets mtctx->jobReady when it prepared a job and no worker is free. ZSTDMT_releaseAllJobResources() erases every job description and keeps the flag. The next frame skips the preparation and posts the erased job. The worker calls ZSTDMT_getCCtx(NULL) (zstdmt_compress.c:697).
  • zstd reaches this state without a reset from the caller. After a failed job it erases the table and starts a new session by itself. Two scripts that drive _handle and call no reset() exit with SIGSEGV on main in 30 of 30 and 29 of 30 runs, and in 0 of 30 with the patch.
  • A session reset in an open frame is one more way in. Node v26.10.0 exits with SIGSEGV on the worker case of the new test in 200 of 200 runs. A debug build of this PR without the patch reports SEGV ... in ZSTDMT_compressionJob zstdmt_compress.c:697 in 5 of 5 runs. With the patch the case passes.
  • The old reset() did not reach it, because it made a new context.
  • facebook/zstd has no issue or PR for it (checked 2026-09-30), and I did not open one. This C program has no bun in it. It gets SIGSEGV in 20 of 20 runs on v1.5.7 (the vendored commit) and on dev at 01b7154f (2026-09-18), and in 0 of 20 with the line of the patch. On dev, a ZSTD_e_continue call after the frame started to end (stage_wrong) and a ZSTD_CCtx_reset(cctx, ZSTD_reset_session_only) in an open frame give the same two results. An open pull request there (number 4805) adds a NULL check to ZSTDMT_getCCtx() and does not clear jobReady.
/* cc -O1 -g -DZSTD_MULTITHREAD -pthread -Ilib -Ilib/common jobready.c \
 *    lib/common/*.c lib/compress/*.c lib/decompress/*.c -DZSTD_DISABLE_ASM
 *
 * The caller never calls ZSTD_CCtx_reset(). */
#include <stdio.h>
#include <stdlib.h>
#include "zstd.h"

int main(void) {
    size_t const job = 1 << 20, cap = ZSTD_compressBound(2 * job);
    char* const src = malloc(2 * job);
    char* const dst = malloc(cap);
    ZSTD_CCtx* const cctx = ZSTD_createCCtx();
    unsigned seed = 1;
    size_t i;
    int call;
    setvbuf(stdout, NULL, _IONBF, 0);
    for (i = 0; i < 2 * job; i++) {
        seed = seed * 1103515245u + 12345u;
        src[i] = "0123456789abcdef"[(seed >> 16) & 15];
    }
    ZSTD_CCtx_setParameter(cctx, ZSTD_c_nbWorkers, 1);
    ZSTD_CCtx_setParameter(cctx, ZSTD_c_jobSize, (int)job);
    /* The pledged size is less than the first job, so the first job fails. */
    ZSTD_CCtx_setPledgedSrcSize(cctx, 600000);
    for (call = 0; call < 3; call++) {
        /* Call 0: job 0 goes to the worker.
         * Call 1: job 1 is prepared, the worker is busy: jobReady = 1.
         * Call 2: reports the error of job 0. zstd resets the session. */
        ZSTD_inBuffer in = { src + (call == 1 ? job : 0), call < 2 ? job : 0, 0 };
        ZSTD_outBuffer out = { dst, cap, 0 };
        size_t const r = ZSTD_compressStream2(cctx, &out, &in, call < 2 ? ZSTD_e_continue : ZSTD_e_end);
        printf("frame 1, call %d: %s\n", call, ZSTD_isError(r) ? ZSTD_getErrorName(r) : "ok");
    }
    {   /* The next frame on the same context posts the erased job. */
        ZSTD_inBuffer in = { src, job + 1, 0 };
        ZSTD_outBuffer out = { dst, cap, 0 };
        size_t const r = ZSTD_compressStream2(cctx, &out, &in, ZSTD_e_end);
        printf("frame 2: %s\n", ZSTD_isError(r) ? ZSTD_getErrorName(r) : "ok");
    }
    ZSTD_freeCCtx(cctx);
    free(src);
    free(dst);
    return 0;
}

Tests

  • test/js/node/zlib/zlib-zstd-reset.test.ts uses node:test. 12 stream tests run in Bun and in Node. The last test runs the file under the node of the machine when that Node is v26.10.0 or later, and 12 of 12 pass there. The test is skipped for an older Node: its reset() drops the options, Node v24.3.0 has no dictionary option for zstd, and Node v20.19.0 cannot load a .ts file. CI has Node v26.3.0, so CI skips this one test.
  • One test starts a second process for 5 handle cases: reset() of a handle with no context (2), reset() while a job waits for the worker, a failed job while a job waits (no reset()), and reset() while the worker reads the dictionary. The worker case checks that the job did wait (the two writes gave no output) and tries again if it did not. After 10 tries it runs in either state, so the scheduler cannot fail it. At load average 710, 1 of 500 first tries missed the state on a release build, and 0 of 300 on a debug build.
  • test/js/node/test/parallel/test-zlib-zstd-reset.js is the file of Node v26.10.0, byte for byte (blob 839669bc63).
  • main (542f52b): 11 of 14 and 3 of 3 fail. With the source of main and the zstd patch (debug, ASAN): the same 14 fail, and the dictionary case reports heap-use-after-free in ZSTDMT_compressionJob with the free in Context::init.
  • Also run on this PR: test/js/node/zlib/ (all files), test/js/web/streams/compression.test.ts, test/js/bun/util/zstd.test.ts, test/js/web/fetch/fetch-compress.test.ts, and the 71 vendored test-zlib*, test-webstreams-*compression* and test-stream-iter-transform* files.

Measurements (release builds, main 542f52b against this PR, linux x64)

  • reset(), gdb breakpoints, 1000 resets of a compressor and 1000 of a decompressor: context frees 2000 to 0, allocations 2000 (101,264,000 B requested) to 0.
  • A used stream, reset() and the next frame: reset() makes 3 allocator calls on main (5,288 B for a compressor, 95,976 B for a decompressor) and 0 now. The next frame allocates 856,217 B (compressor) or 131,072 B (decompressor) again on main and nothing now.
  • Never-reset path: valgrind, perf and strace are not on this machine, so there is no instruction count. The instruction sequences of NativeZstdPrototype__write (729), __writeSync (568), NativeZstdClass__construct (227), ZSTD_compressStream2 (544) and ZSTD_decompressStream (787) are the same in the two binaries after addresses are removed. Context has no new field.
  • zstd patch: ZSTDMT_releaseAllJobResources 99 to 106 instructions (1 store, 6 of padding), 394 to 404 bytes. 20,000 one-shot zstdCompressSync calls with default options call it 0 times in the two builds.
  • Binary: text 80,666,326 to 80,666,134 bytes. The file size is the same (80,832,072). 3 of 78,102 functions change size: NativeZstdPrototype__reset 223 to 261 bytes, NativeZstdPrototype__init 2112 to 2884 (it now contains Context::init, 867 bytes, which has one caller left), ZSTDMT_releaseAllJobResources. bloaty is not installed. The numbers are from size and nm -S.
  • Memory that a used stream keeps from reset() to close() (ZSTD_sizeof_CCtx and ZSTD_sizeof_DCtx, a harness linked to the zstd objects of the release build): default compressor 3,663,393 B, level 19 93,848,223 B, nbWorkers 2 51,680,744 B and 2 idle threads, decompressor after a windowLog 27 frame 134,706,984 B. main: 5,288 B and 95,976 B, because it made a new context. A stream that is used again allocates these buffers again on main, so the peak is the same. bun reports 5,272 B and 95,968 B to the GC in the two builds.
  • Workers, level 19, 2 jobs of 1 MiB in flight, 5 interleaved runs: reset() blocks the JS thread for 829 to 1147 ms on main (median 960) and 0 ms now. The first write after it takes 0.13 to 0.45 ms on main and 574 to 983 ms now (median 616). zstd waits for the running jobs in the two cases. For a stream write that wait is on the thread pool. 100 cycles of jobs and reset() make 1 worker pool and free 1 in the two builds.
  • The new test file on the debug ASAN build: 6.9 s to 13.0 s for 14 tests in 5 runs, on a machine with a load average above 300.

Behaviour that changes outside the report

Open pull request on the same struct

#33413 adds decoder state to Context (decode, frame_prefix_size, possible_frame_types) and clears it in init(). The two branches merge with no text conflict. reset() here names every field of Context in a pattern with no .., so the merged tree does not compile until reset() clears the new fields, as Node's ZstdDecompressContext::ResetStream does.

Not in this PR

reset() of a zstd stream freed the ZSTD_CCtx or ZSTD_DCtx and made a new one,
so the stream lost its dictionary and its parameters. It now resets the
session on the same context and sets the pledged size again, as Node does
since v26.10.0 (nodejs/node#65867).

A zstd patch clears mtctx->jobReady where zstd erases its job table. With
ZSTD_c_nbWorkers, zstd kept that flag after it erased a job that waited for a
worker, and the next frame on the context ran the erased job (SIGSEGV). zstd
reaches that state by itself after a job fails. A session reset in an open
frame reaches it too, so reset() needs the patch.
@robobun

robobun commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Status

Reproduced on bun 1.4.3 canary (367d939) and on a release build of main (542f52b) with the script in the description. After reset(), a zstd decompressor with a dictionary fails a valid frame with ZSTD_error_corruption_detected. Node v26.10.0 decodes the frame. Node v26.3.0 fails as main does.

Pull request: #44261

@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: oven-sh/bun/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: a5fd8f80-0c88-4446-aa4b-6030ff1f207d

📥 Commits

Reviewing files that changed from the base of the PR and between 1b18197 and 6bf1d1a.

📒 Files selected for processing (1)
  • test/js/node/zlib/zlib-zstd-reset.test.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 6 remain after this review.


Walkthrough

Context::reset now performs a session-only Zstandard reset and preserves the existing context’s dictionary and parameters. Compressor resets reapply the pledged source size. A dependency patch clears stale multithreaded job state, and new tests cover reset behavior.

Changes

Zstandard stream reset

Layer / File(s) Summary
Session reset and job cleanup
src/zstd/lib.rs, src/runtime/node/zlib/NativeZstd.rs, patches/zstd/mt-clear-job-ready.patch, scripts/build/deps/zstd.ts, src/runtime/node/node_zlib_binding.rs
Context::reset uses a session-only reset and preserves the existing context’s dictionary and parameters. Compressor resets reapply the pledged source size. The zstd patch clears jobReady after job resources are erased, and the dependency configuration applies that patch. The reset comment clarifies affected native state.
Reset behavior tests
test/js/node/test/parallel/test-zlib-zstd-reset.js, test/js/node/zlib/zlib-zstd-reset.test.ts
Tests cover compressor and decompressor resets, retained dictionaries and parameters, pledged-size validation, window limits, and multithreaded job cases.

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 6bf1d

The supplied change only adds tests for the zstd reset behavior, and no merge-blocking issue was found in them.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely identifies the primary change: resetting Zstd streams in place while preserving dictionaries and parameters.
Description check ✅ Passed The description is comprehensive and explains the problem, implementation, verification, risks, and scope. It does not use the exact template headings, but it includes the required information in equi…

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 8:35 PM PT - Sep 29th, 2026

✅ @robobun, your commit 6bf1d1af55332bc1ad21a8e9e30741ab02a20190 passed in Build #121758! 🎉


🧪   To try this PR locally:

bunx bun-pr 44261

That installs a local version of the PR into your bun-44261 executable, so you can run:

bun-44261 --bun

Node v24.3.0 has no dictionary option for zstd streams. The test of the
options failed when `node` on the machine was that version. An older
Node now skips all 12 stream tests.
Comment thread src/runtime/node/node_zlib_binding.rs Outdated
Comment thread src/runtime/node/zlib/NativeZstd.rs Outdated
Comment thread src/runtime/node/zlib/NativeZstd.rs Outdated
Comment thread src/runtime/node/zlib/NativeZstd.rs Outdated
Each comment that reset() adds is one line. The comment in
CompressionStream::reset keeps the text of main and changes one line.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🟣 src/runtime/node/zlib/NativeZstd.rs — A process that writes to a multithreaded zstd compressor with a dictionary and then close()s it can crash with a heap use-after-free in the worker thread, with or without a reset(). close() at NativeZstd.rs:698 calls ZSTD_CCtx_reset(session_and_parameters), whose parameter half runs ZSTD_clearAllDicts and frees the cdict while a ZSTD_c_nbWorkers job that reads it may still be running; the worker is only joined later inside ZSTD_freeCCtx at NativeZstd.rs:712. The PR touches the reset/close lifecycle and its own fixture comment names this order (#44201) but leaves the close() site as is. …

    Why this was flagged

    …Fix: join outstanding jobs before dropping the dictionary on every teardown path, e.g. call ZSTD_freeCCtx directly (it joins the pool before freeing) or a session-only reset followed by a flush that waits for the workers, and apply the same to the reset()+close() sequence.

    Trigger: createZstdCompress({ dictionary, params: { [ZSTD_c_nbWorkers]: 1 } }), one write() large enough to post a job (>= 512 KiB or ZSTD_c_jobSize), then close() or stream destruction while the worker is still compressing; reached through CompressionStream::close -> Context::close at NativeZstd.rs:688. At NativeZstd.rs:698-701 close() calls ZSTD_CCtx_reset(state, ZSTD_reset_session_and_parameters): the session half sets streamStage to init, so the parameters half passes its stage check and runs ZSTD_clearAllDicts, freeing the cdict the running ZSTDMT_compressionJob dereferences. Only afterwards does deinit_state at NativeZstd.rs:712 call ZSTD_freeCCtx, which joins the pool. The base branch has the same close() body, so this is a pre-existing crash, but the diff makes reset() no longer free the context (the old…

    Verification: pre-existing. Trigger: a zstd compressor created with dictionary and ZSTD_c_nbWorkers >= 1, one write large enough to post a job (>= jobSize; default target job size at level 3 is 8 MiB, 512 KiB when ZSTD_c_jobSize is set, or any filled input on ZSTD_e_flush), then close()/destroy() before the frame ends and while the worker is still on the first job. Mechanism: Context::close at…

Comment thread src/runtime/node/zlib/NativeZstd.rs Outdated
Comment thread src/runtime/node/zlib/NativeZstd.rs
Comment thread test/js/node/zlib/zlib-zstd-reset.test.ts Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Still open from earlier reviews (3):

  • Unresolved: 3 minor or pre-existing.

Comment thread test/js/node/zlib/zlib-zstd-reset.test.ts
…d Node

The case "reset() while a job waits for the worker" still tries up to 10
times to reach that state. It no longer fails when the last try does not
reach it.

The file runs under Node.js only when the `node` of the machine is
v26.10.0 or later. An older Node drops the options in reset(), and
Node 20 cannot load a .ts file. The test is skipped there.
@robobun

robobun commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

The close() finding in the review above is #44201. main has it with no reset(), and this PR does not change close().

The old reset() was a fourth place with the same order, and this PR removes it. The other three places are close(), the finalizer and a second init(). A direct ZSTD_freeCCtx does not fix them: #44201 shows that it clears the dictionary before it joins the workers.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Comment on lines +49 to +53
// With nbWorkers >= 1, zstd keeps its jobReady flag when it erases the
// job table, so the next frame on that context posts an erased job to a
// worker (SIGSEGV). zstd gets there by itself after a failed job, and
// node:zlib reset() gets there too. Not reported upstream yet.
"patches/zstd/mt-clear-job-ready.patch",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 nit (optional): Maintainers get a vendored zstd bug-fix patch with no upstream issue to track and no self-obsoleting check, so it can outlive the upstream fix silently. scripts/build/deps/zstd.ts:53 registers mt-clear-job-ready.patch with only a comment saying "Not reported upstream yet", and nothing is added to scripts/build/workarounds.ts. Fix: file the upstream facebook/zstd issue (the PR already has a standalone C repro) and link it from the patch header and this comment, and add a workarounds.ts entry whose expectedToBeFixed trips when ZSTD_COMMIT moves past the pinned commit, with a cleanup string naming the patch to drop.

Why this was flagged

The diff adds patches/zstd/mt-clear-job-ready.patch and lists it in the patches array at scripts/build/deps/zstd.ts:49-53; the patch header (patches/zstd/mt-clear-job-ready.patch:22-23) says no upstream issue exists and the zstd.ts comment says "Not reported upstream yet". scripts/build/CLAUDE.md "Adding a workaround" says every temporary fix waiting on an upstream release registers an entry in scripts/build/workarounds.ts with an expectedToBeFixed predicate, and workarounds.ts:9-11 explicitly lists "vendored dep bump" as such a case; the registry currently has only two entries (workarounds.ts:67-119), none for this patch. The consequence is operational for maintainers: when ZSTD_COMMIT is bumped later, the one-line hunk will still apply cleanly on top of an upstream fix (or a refactor that moved the bug), so nothing tells the developer to re-evaluate or drop the patch, and there is no upstream reference to check against. On the base branch this patch and this class of tracking gap do not exist.

Verification: nit. Triggering condition: a future zstd bump where upstream fixes jobReady differently (so the Bun hunk still applies) — nothing then tells anyone the patch is obsolete. Verified facts: (1) scripts/build/deps/zstd.ts:49-53 (diff) adds "patches/zstd/mt-clear-job-ready.patch" with the comment "... Not reported upstream yet."; (2) the patch header…

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No code change for this one.

The upstream issue. There is none yet, and I did not open one on facebook/zstd. The Notes in the description have what a report needs: a C program with no bun in it, its result on v1.5.7 and on dev at 01b7154f (SIGSEGV in 20 of 20 runs), and its result with the one line of the patch (0 of 20). When an issue exists, its link belongs in the header of the patch and in the comment here.

The workarounds.ts entry. None of the other 30 patches in patches/ has one. The two entries there are for the toolchain and for the libc crate. A check on ZSTD_COMMIT stops every zstd update at configure, also when upstream has no fix. The test in this PR covers the other direction: without the patch, the handle case of zlib-zstd-reset.test.ts crashes in ZSTDMT_compressionJob while the bug is there. If a maintainer wants the entry, I will add it.

@robobun

robobun commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

A note for the merge with #44421, the fix for #44201. Both PRs change how a zstd handle is torn down. Measured on a debug + ASAN build with both diffs applied (head 6bf1d1a of this PR, 2026-10-01):

// in the block "zstd: a compressor with workers and a dictionary"
[
  "reset(), a write and destroy()",
  `const stream = open();
   await write(stream, input.subarray(0, 1000));
   stream.reset();
   await write(stream, input);
   stream.destroy();
   console.log("destroyed");`,
  "destroyed",
],

Neither PR needs the other to be correct by itself. The row above was measured with an earlier state of the block (a 123-byte dictionary).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant