Skip to content

bun test: detect obsolete snapshots and fix -u "added" mislabel - #34042

Closed
robobun wants to merge 11 commits into
mainfrom
farm/f88d2771/snapshot-obsolete-detection
Closed

robobun wants to merge 11 commits into
mainfrom
farm/f88d2771/snapshot-obsolete-detection

Conversation

@robobun

@robobun robobun commented Jul 12, 2026 •

Copy link
Copy Markdown
Collaborator

What

bun test now reports snapshot keys that exist in a .snap file but were never matched during the run, and bun test -u stops mislabeling rewritten entries as "added".

Fixes #12114

Problem

$ bun test           # .snap file has a stale `obsolete-gone 1` key
 1 pass
 0 fail
 1 snapshots, 1 expect() calls     # no signal that a stale key exists

$ bun test -u
snapshots: +1 added                # nothing was added; one entry was rewritten, one removed

Jest prints "1 snapshot obsolete" on the first run and "1 snapshot removed" on the second. Bun had no way to represent "in the file but never matched": Snapshots carried only added/passed/failed, and under -u the file was opened with O_TRUNC so every entry looked brand new. The early truncate also left the .snap file empty for anything that read it during afterAll() (#12114).

Fix

src/runtime/test_runner/snapshot.rs

  • Snapshots gains obsolete, removed counters and an unchecked_keys: StringHashMap<()> populated by parse_file and drained on each match in get_or_put.
  • Under -u, get_snapshot_file reads and parses the old file before clearing the buffer so a key that already existed counts as passed rather than added; write_snapshot_file truncates after write instead of relying on O_TRUNC at open. A malformed old file is ignored under -u so the existing "replace unparseable file" behaviour is preserved.
  • mark_snapshots_as_checked_for_test removes every {testName} N key for a given test (Jest's keyToTestName rule) so skipped/todo/filtered tests are not counted as obsolete.

src/runtime/test_runner/Execution.rs

  • on_sequence_completed calls the new hook for Skip / Todo / SkippedBecauseLabel results, building the name identically to Expect::get_snapshot_name.

src/runtime/cli/test_command.rs

  • Summary prints N removed / N obsolete (with a bun test -u hint). The obsolete line is suppressed under --only since non-only tests never reach the reporter there.

Verification

$ bun test
snapshots: 1 passed, 1 obsolete
To remove obsolete snapshots, run bun test -u

$ bun test -u
snapshots: 1 passed, 1 removed

New coverage in test/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts: plain-run obsolete count, -u removed vs added, test.skip and -t not producing false positives, a truly new key still reporting added, and the #12114 repro (.snap file readable in afterAll under -u). All six fail on current main and pass with this change. Existing snapshot-tests/, ci-restrictions.test.ts, and bun_test.test.ts pass unchanged.


[review] gate passed · iteration 3 · 6 files touched

fails on main (without fix)
ASAN without fix: 9 FAILED
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts
info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu
info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05)
info: component rust-src is up to date
info: checking for self-update (current version: 1.29.0)
bun test v1.4.0 (a024016cd)

test/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts:
31 |         "// Bun Snapshot v1, https://bun.sh/docs/test/snapshots\n\n" +
32 |         "exports[`foo 1`] = `1`;\n\n" +
33 |         "exports[`foo 2`] = `2`;\n",
34 |     });
35 |     const { stderr, exitCode } = await run(dir);
36 |     expect(stderr).toContain("1 obsolete");
                        ^
error: expect(received).toContain(expected)

Expected to contain: "1 obsolete"
Received: "\nsnap.test.ts:\n(pass) foo [3.09ms]\n\n 1 pass\n 0 fail\n 1 snapshots, 1 expect() calls\nRan 1 test across 1 file. [395.00ms]\n"

      at <anonymous> (/workspace/bun/test/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts:36:20)
(fail) obsolete 
... (truncated)

release without fix: all passed
bun test v1.4.0-canary.1 (d9ac81530)

test/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts:
(pass) obsolete snapshot detection > passing test with fewer toMatchSnapshot calls than keys on disk reports obsolete [17.65ms]
(pass) obsolete snapshot detection > -u labels dropped entries as removed, not added [20.44ms]
(pass) obsolete snapshot detection > reports obsolete snapshots without -u [21.28ms]
(pass) obsolete snapshot detection > skipped test's snapshot is not counted obsolete (skip before first match) [18.43ms]
(pass) obsolete snapshot detection > -u counts a skipped test's dropped entry in removed [18.30ms]
(pass) obsolete snapshot detection > in-source test.only() suppresses obsolete [17.42ms]
(pass) obsolete snapshot detection > skipped test's snapshot is not counted obsolete (skip after first match) [20.23ms]
(pass) obsolete snapshot detection > beforeAll failure suppresses obsolete for jumped-over tests [16.77ms]
(pass) obsolete snapshot detection > skipped test with a hinted snapshot is not counted obsolete [15.67ms]
(pass) obsolete snapshot detection > -t filtered test's snapshot is not counted obsolete [14.25ms]
(pass) obsolete snapshot detection >
... (truncated)
passes on PR (with fix)
ASAN with fix: all passed
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts
info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu
info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05)
info: component rust-src is up to date
info: checking for self-update (current version: 1.29.0)
bun test v1.4.0 (a024016cd)

test/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts:
(pass) obsolete snapshot detection > passing test with fewer toMatchSnapshot calls than keys on disk reports obsolete [486.97ms]
(pass) obsolete snapshot detection > reports obsolete snapshots without -u [474.96ms]
(pass) obsolete snapshot detection > skipped test's snapshot is not counted obsolete (skip after first match) [489.76ms]
(pass) obsolete snapshot detection > skipped test's snapshot is not counted obsolete (skip before first match) [507.33ms]
(pass) obsolete snapshot detection > -u labels dropped entries as removed, not added [543.50ms]
(pass) obsolete snapshot detection > -u counts a skipped test's dropped entry in removed [4
... (truncated)

release with fix: all passed
$ bun scripts/build.ts --profile=release
info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu
info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05)
info: component rust-src is up to date
info: checking for self-update (current version: 1.29.0)
[configured] bun-profile → bun (stripped) in 675ms (unchanged)
ninja: Entering directory `/workspace/bun/build/release'
[1/6] gen generated_host_exports.rs
generated_host_exports.rs: 91 exports (host=3, lazy=10, generic=78, rust=0); 243 extern-C blocks audited
[1/6] cargo bun_bin → libbun_rust.a (--target x86_64-unknown-linux-gnu)
info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu
info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05)
info: component rust-src is up to date
info: component rust-std is up to date

  nightly-2026-05-06-x86_64-unknown-linux-gnu unchanged - rustc 1.97.0-nightly (e95e73209 2026-05-05)

info: checking for self-update (current version: 1.29.0)
�[1m�[92m   Compiling�[0m bun_core v0.0.0 (/workspace/bun/src/bun_core)
�[1m�[92m   Compiling�[0m bun_errno v0.0.0 (/workspace/bun/src/errno)
�[1m�
... (truncated)
diff hotspot
src/runtime/cli/test_command.rs                    |  45 ++++-
 src/runtime/test_runner/Execution.rs               |  22 +++
 src/runtime/test_runner/bun_test.rs                |   4 +-
 src/runtime/test_runner/expect.rs                  |  14 +-
 src/runtime/test_runner/snapshot.rs                | 137 +++++++++++--
 .../test/snapshot-tests/obsolete-snapshots.test.ts | 212 +++++++++++++++++++++
 6 files changed, 408 insertions(+), 26 deletions(-)

gate history · 7 passed · 0 rejected · iteration 3

evidence per changed file
file                                                      reads  edits  tests
src/runtime/cli/test_command.rs                               8     17      0
src/runtime/test_runner/Execution.rs                         11     14      0
src/runtime/test_runner/bun_test.rs                           2      1      0
src/runtime/test_runner/expect.rs                             3      1      0
src/runtime/test_runner/snapshot.rs                          13     26      0
…t/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts      1     10      0

…as added

Snapshots loaded from the .snap file but never matched during a run were
previously invisible: a plain run gave no signal and -u silently rebuilt
the file while booking every rewritten entry through the "added" counter.

Snapshots now tracks an unchecked_keys set populated at parse time and
drained on each match. Remaining entries at file close are tallied as
obsolete (plain run) or removed (-u). Under -u the old file is read
before the rebuild so a key that already existed counts as passed rather
than added, and the file is truncated after write instead of at open.
Skipped, todo, and -t filtered tests mark their keys checked so they are
not reported as obsolete; --only suppresses the obsolete line entirely
since those tests never reach the reporter.
@robobun

robobun commented Jul 12, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 11:09 AM PT - Jul 12th, 2026

❌ @robobun, your commit a024016 has 1 failures in Build #72303 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 34042

That installs a local version of the PR into your bun-34042 executable, so you can run:

bun-34042 --bun

snapshots/snapshot.test.ts has an env-dependent pre-existing failure
(error snapshots needs FORCE_COLOR) that would mask the fail-before
check; keep the new coverage in its own file.
@github-actions

Copy link
Copy Markdown
Contributor

Found 1 issue this PR may fix:

  1. On --update-snapshots, snapshot file remains empty in afterAll() #12114 - PR fixes the O_TRUNC behavior that caused snapshot files to be empty during afterAll() when using --update-snapshots; the file is now read and parsed before being rewritten

If this is helpful, copy the block below into the PR description to auto-close this issue on merge.

Fixes #12114

🤖 Generated with Claude Code

Comment thread src/runtime/test_runner/Execution.rs Outdated
Comment thread src/runtime/test_runner/snapshot.rs Outdated
@coderabbitai

coderabbitai Bot commented Jul 12, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 7 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 1c99f5e9-f472-40c9-bcc5-fcda760b5069

📥 Commits

Reviewing files that changed from the base of the PR and between d295900 and 703e806.

📒 Files selected for processing (4)
  • src/runtime/cli/test_command.rs
  • src/runtime/test_runner/Execution.rs
  • src/runtime/test_runner/snapshot.rs
  • test/js/bun/test/snapshot-tests/obsolete-snapshots.test.ts

Walkthrough

Changes

The snapshot runner now tracks unchecked snapshot keys to classify matched, obsolete, and removed entries. Skipped tests are marked checked, CLI output reports new counters, and snapshot tests cover update, filtering, and skip behavior.

Snapshot obsolete tracking

Layer / File(s) Summary
Snapshot state and accounting
src/runtime/test_runner/snapshot.rs, src/runtime/cli/test_command.rs
Snapshot keys are loaded, matched, and classified as passed, added, obsolete, or removed across normal and update modes.
Execution snapshot filtering
src/runtime/test_runner/Execution.rs
Skipped, todo, and label-skipped tests mark their snapshot entries as checked using Jest-compatible names.
CLI reporting and validation
src/runtime/cli/test_command.rs, test/js/bun/test/snapshot-tests/snapshots/snapshot.test.ts
CLI summaries include obsolete and removed counts, with tests covering updates, filters, skipped tests, and combined additions/removals.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title is concise and accurately summarizes the main snapshot reporting change.
Description check ✅ Passed The description covers the PR purpose, problem, fix, and verification, though it uses different headings than the template.

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/runtime/test_runner/Execution.rs`:
- Around line 711-722: Update snapshot bookkeeping for tests excluded by .only()
in addition to the existing skipped, todo, and label-filtered ExecutionSequence
results. Ensure each filtered-out sibling’s snapshot name is passed to Jest’s
mark_snapshots_as_checked_for_test before snapshot cleanup, even when no
ExecutionSequence is created, while preserving the existing handling for
sequences that do exist.

In `@src/runtime/test_runner/snapshot.rs`:
- Around line 359-362: Handle the return value of bun_sys::ftruncate in the
update_snapshots branch instead of discarding it. Propagate the failure using
the same fatal error behavior as write_inline_snapshots, ensuring the snapshot
rewrite cannot return success while stale trailing data remains; keep the
subsequent file.file.close flow unchanged for successful truncation.

In `@test/js/bun/test/snapshot-tests/snapshots/snapshot.test.ts`:
- Around line 957-1026: Add a snapshot test alongside the obsolete/update cases
that combines test.only(...) with the -u option, including a selected test and a
sibling test with existing snapshots. Assert the run succeeds and the rewritten
snapshot file still contains the sibling snapshot, while preserving the selected
test’s update behavior. Use the existing temp-directory, snapFixture, run, and
snapshot-file assertions.
- Around line 936-1027: Make the “obsolete snapshot detection” suite concurrent
by changing its describe declaration to describe.concurrent, allowing all
contained subprocess and file-I/O tests to run in parallel. Preserve the
existing test bodies, fixtures, assertions, and isolation behavior.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: c86e2eef-6f76-4ef8-b3ff-45304cdb9cb9

📥 Commits

Reviewing files that changed from the base of the PR and between 2e2230a and d295900.

📒 Files selected for processing (4)
  • src/runtime/cli/test_command.rs
  • src/runtime/test_runner/Execution.rs
  • src/runtime/test_runner/snapshot.rs
  • test/js/bun/test/snapshot-tests/snapshots/snapshot.test.ts

Comment thread src/runtime/test_runner/Execution.rs Outdated
Comment thread src/runtime/test_runner/snapshot.rs
Comment thread test/js/bun/test/snapshot-tests/snapshots/snapshot.test.ts Outdated
Comment thread test/js/bun/test/snapshot-tests/snapshots/snapshot.test.ts Outdated
…ftruncate error

Skipped/todo/filtered test names are now recorded with their file id and
reconciled against unchecked_keys in write_snapshot_file, so a test.skip
declared before the first toMatchSnapshot call is still excluded from the
obsolete count. Under -u the reconciliation is skipped so the removed
tally matches what is physically dropped from the .snap file. ftruncate
failure after the -u rewrite now propagates as FailedToWriteSnapshotFile
instead of returning Ok with stale trailing bytes.

Tests: skip-before/after-first-match ordering, -u + skip removal tally,
describe made concurrent.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional findings (outside current diff — PR may have been updated during review):

  • 🔴 src/runtime/cli/test_command.rs:2855-2859 — The obsolete-count suppression if reporter.jest.only { 0 } only covers the CLI --only flag — in-source test.only() sets scope.base.only = Only::Contains on the describe tree and never touches reporter.jest.only, and non-only entries are dropped at Order.rs:103 before ever getting an ExecutionSequence, so their snapshot keys stay in unchecked_keys and are reported as obsolete. Every developer who temporarily focuses a test with .only() will see a spurious "N obsolete" plus a hint to run bun test -u, which then physically deletes the non-only tests' still-valid snapshots. Fix by also suppressing when the file's root scope ended up Only::Contains (or set a "had in-source only" flag on the runner/snapshots that the summary check reads).

    Extended reasoning...

    What the bug is

    The PR suppresses the obsolete-snapshot count when --only is in effect, because non-only tests never run and their snapshot keys would otherwise all show up as obsolete:

    let obsolete = if reporter.jest.only { 0 } else { reporter.jest.snapshots.obsolete };

    But reporter.jest.only (TestRunner.only, jest.rs:116) reflects only the CLI --only flag. It is initialized from ctx.test_options.only and the sole write to it anywhere in src/runtime/ is the scopeguard restore at test_command.rs:3121, which puts it back to the CLI value after each module. In-source test.only() goes through a completely separate path: BaseScopeCfg { self_only: true } → base.only = Only::Yes (bun_test.rs:1701) → mark_contains_only() walks the parent chain setting scope.base.only = Only::Contains (bun_test.rs:1758–1769). Nothing on that path ever sets runner.only.

    Why the mark-as-checked hook doesn't help here

    The PR's new mark_snapshots_as_checked_for_test hook fires from on_sequence_completed for Skip / Todo / SkippedBecauseLabel results. But when a file uses in-source .only(), non-only entries never get that far: Order::generate_order reads scope_only = current.base.only and, when it is Only::Contains, continues straight past every entry whose only == Only::No (Order.rs:101–105). Those entries are never handed to generate_order_sub, so no ExecutionSequence is created for them, on_sequence_completed never fires, and their snapshot keys are never removed from unchecked_keys.

    This is a different mechanism from the already-reported ordering bug on this PR: that one is about skipped/filtered tests that are sequenced but complete before the .snap file is lazily opened. This one is about tests that are never sequenced at all, so no amount of deferring the mark-as-checked call will reach them.

    It also differs from -t filtering: -t sets ScopeMode::FilteredOut, which still creates a sequence and completes with Result::SkippedBecauseLabel, so the hook fires. .only() drops entries at order-generation time.

    Step-by-step proof

    Given snap.test.ts:

    test.only("a", () => expect(1).toMatchSnapshot());
    test("b", () => expect(2).toMatchSnapshot());

    and __snapshots__/snap.test.ts.snap containing both `a 1` and `b 1`, running bun test (no --only flag):

    1. Collection: test.only("a", …) sets its own only = Only::Yes and calls mark_contains_only() on its parent chain → root describe's base.only = Only::Contains. test("b", …) has only = Only::No. runner.only stays false (CLI flag not passed).
    2. Order::generate_order on the root scope: scope_only == Only::Contains, so at Order.rs:103 the entry for b (only == Only::No) hits continue — no sequence is generated for it.
    3. Execution: only a runs. toMatchSnapshot() → get_snapshot_file opens the .snap, parse_file populates unchecked_keys = {"a 1", "b 1"}. get_or_put removes "a 1". "b 1" remains.
    4. No sequence for b exists, so on_sequence_completed never fires for it and mark_snapshots_as_checked_for_test("b") is never called.
    5. write_snapshot_file: unchecked_keys.len() == 1 → self.obsolete += 1.
    6. Summary at test_command.rs:2855: reporter.jest.only is false (it was never set — and even if something per-module had set it, the scopeguard at 3121 restored it to the CLI value before this point). The suppression does not fire. Output: snapshots: 1 passed, 1 obsolete plus To remove obsolete snapshots, run bun test -u.

    Following that hint runs the same order-generation drop, b never writes into the rebuilt file_buf, and write_snapshot_file truncates — `b 1` is physically deleted from disk and reported as 1 removed.

    Impact

    test.only() is one of the most common iteration workflows — temporarily focusing on one test while editing. With this PR every such run prints a false-positive obsolete warning for every other snapshot-using test in the file, and the accompanying hint steers the user toward -u, which then destroys those tests' snapshots. When .only() is later removed, all of them come back as added with no diff against their previous stored values. The PR author clearly intended to suppress this case (the --only guard is there for exactly this reason) but only wired it to the CLI flag.

    The new test file covers test.skip and -t but not in-source test.only(), so this is unexercised.

    How to fix

    The cleanest option is to suppress per-file at the point the count is produced: when a file's root describe ends up Only::Contains because of in-source .only(), skip the self.obsolete += unchecked tally in write_snapshot_file for that file (e.g. set a had_in_source_only: bool on Snapshots from mark_contains_only, or check the root scope's only state before calling write_snapshot_file). Alternatively, keep the summary-time check but read a flag that is set whenever any file used in-source .only(), not just the CLI flag.

    Add a test alongside the existing test.skip / -t cases that uses test.only("keeps", …) + test("skipped", …) and asserts the summary does not contain obsolete, plus a -u variant asserting skipped 1 survives on disk.

Comment thread src/runtime/test_runner/Execution.rs Outdated
Comment thread src/runtime/test_runner/snapshot.rs
note_skipped_test is now called for every non-Pass sequence result (Jest
calls markSnapshotsAsCheckedForTest on test_done for pass and fail), so a
test that throws before reaching its toMatchSnapshot is not flagged
obsolete. The root describe's Only state is checked per file so both CLI
--only and in-source test.only() suppress the obsolete tally for that
file, replacing the summary-time reporter.jest.only check which only
covered the CLI flag.
@robobun

robobun commented Jul 12, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the additional findings in 703e806:

  • in-source test.only(): the per-file obsolete tally is now skipped when the file's root describe has Only::Contains set (covers both CLI --only and in-source test.only()), replacing the summary-time reporter.jest.only check which only covered the CLI flag. Test added.
  • Result::Fail / FailBecause*: note_skipped_test is now called for every non-Pass result, so a test that throws before its toMatchSnapshot is not flagged obsolete. Test added.
  • ftruncate error: already propagated as FailedToWriteSnapshotFile in a2e25d0.

10 tests in obsolete-snapshots.test.ts; 8 fail on canary, all 10 pass on this branch.

Comment thread src/runtime/cli/test_command.rs Outdated
Comment thread src/runtime/test_runner/Execution.rs Outdated
Comment thread src/runtime/test_runner/snapshot.rs
Comment thread src/runtime/test_runner/snapshot.rs
…pped-names vec

file_was_partial replaces had_only_in_file and is also set on any failing
sequence (including hook-only sequences), so tests jumped over by a
beforeAll failure or a throwing describe callback do not produce a false
obsolete hint on an already-red run.

mark_snapshots_as_checked_for_test also matches '{test_name}: {hint}'
prefixes so a skipped test with a hinted toMatchSnapshot is not flagged.

get_snapshot_file prunes skipped_test_names to the file being opened so
entries from files with no snapshot calls do not accumulate.
@robobun

robobun commented Jul 12, 2026 •

Copy link
Copy Markdown
Collaborator Author

CI is red on test/js/node/test/parallel/test-worker-message-port-transfer-terminate.js (SIGABRT on x64-asan, JSC getOwnPropertyDescriptor assertion + OOM). That failure is unrelated to this diff, which only touches snapshot.rs / Execution.rs / expect.rs / bun_test.rs / test_command.rs / the new obsolete-snapshots.test.ts; it reproduced identically on five consecutive runs (72264, 72273, 72287, 72299, 72303) and is being tracked separately. The new test file and the existing snapshot / ci-restrictions / bun_test suites are green locally on the debug build. Ready for a maintainer to look past the unrelated asan lane.

Comment thread src/runtime/test_runner/Execution.rs Outdated
Comment thread src/runtime/test_runner/snapshot.rs Outdated
Dropping the '!= Pass' guard means a test.failing() body that throws
(result is Pass with maybe_skip) still has its key excluded from the
obsolete tally; noting a fully-passed test is a no-op since get_or_put
already removed its keys.

partial_file_id replaces the bool so a file with .only() or a failure
that never opens a .snap file cannot suppress the next file's obsolete
tally.
Comment thread src/runtime/test_runner/Execution.rs
A throwing describe callback is handled during collection and never
produces a failing ExecutionSequence, so the per-file partial_file_id
check does not catch it. The run is already red via
unhandled_errors_between_tests, so gate the obsolete hint on that
counter being zero.
Comment thread src/runtime/test_runner/Execution.rs
Comment thread src/runtime/test_runner/Execution.rs Outdated
…cted; share name builder

Noting a cleanly passing test's name caused mark_snapshots_as_checked_for_test
to strip that test's higher-numbered keys (e.g. 'foo 2' when only 'foo 1' was
matched), defeating the headline use case of obsolete detection. Restore the
'!= Pass' guard and add a ScopeMode::Failing arm so a test.failing body that
throws is still noted.

Expect::build_snapshot_name is extracted from get_snapshot_name and reused by
on_sequence_completed so the key format has a single source of truth.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional findings (outside current diff — PR may have been updated during review):

  • 🟡 src/runtime/test_runner/Execution.rs:717-719 — 🟡 partial_file_id is a single Option<FileId> written by plain assignment, so it can be overwritten cross-file: if file A sets Some(A) (via .only() or a fail) and then file B's first sequence completes (fail or .only()) before B's first toMatchSnapshot(), B writes Some(B) while _current_file is still Some(A); A's deferred flush then sees Some(B) != Some(A) and falsely tallies A's .only()-sibling key as obsolete. This is the reverse of the false-negative 8057d85 fixed — B no longer suppresses A, but B can now clobber A's own suppression. Reporting-only and a narrow multi-file ordering, so not blocking; a small Vec<FileId> (or a per-file-boundary write_snapshot_file()) would close it.

    Extended reasoning...

    What the bug is

    Commit 8057d85 replaced file_was_partial: bool with partial_file_id: Option<FileId> so a later file B setting the flag can't suppress an earlier file A's obsolete tally (the false-negative direction). But it's still a single slot written by plain assignment at Execution.rs:718 — runner.snapshots.partial_file_id = Some(file_id) — so the reverse case is now broken: if A wrote Some(A) for its own suppression, B's first completing sequence overwrites it to Some(B), and A's deferred flush loses the suppression it needed.

    Code path

    There is no per-file-boundary write_snapshot_file() in the sequential loop — its three call sites are the lazy switch inside get_snapshot_file (snapshot.rs:~908), end-of-run (test_command.rs:~2645), and parallel-worker teardown (runner.rs:~731, whose comment reads "Snapshots flush lazily when the next file opens its snapshot file"). So file B's on_sequence_completed runs while _current_file is still Some(A). The flush check at snapshot.rs:~361 is else if self.partial_file_id != Some(file.id) — with partial_file_id == Some(B) and file.id == A that's true, so the obsolete branch runs for A.

    Note that skipped_test_names already avoids this class by carrying (FileId, _) per entry and filtering on *id == file.id; partial_file_id has no such per-file protection because it holds only one id.

    Step-by-step proof

    Two files, sequential run, no -u:

    • a.test.ts: test.only("keeps", () => expect({a:1}).toMatchSnapshot()); test("sibling", () => expect({b:2}).toMatchSnapshot()); — __snapshots__/a.test.ts.snap contains both keeps 1 and sibling 1.
    • b.test.ts: test("first", () => { throw new Error("boom") }); test("second", () => expect({c:3}).toMatchSnapshot());
    1. A's keeps runs → get_snapshot_file(A) opens A's .snap, unchecked_keys = {"keeps 1", "sibling 1"}, removes "keeps 1". _current_file = Some(A).
    2. A's keeps completes → on_sequence_completed: root_only == Only::Contains → partial_file_id = Some(A). (sibling was dropped in Order::generate_order before sequencing, so note_skipped_test never fires for it.)
    3. File A finishes; file B starts. No flush — _current_file is still Some(A).
    4. B's first throws → on_sequence_completed: sequence.result.is_fail() → partial_file_id = Some(B) (overwrites A).
    5. B's second calls toMatchSnapshot() → get_snapshot_file(B) → write_snapshot_file() for A. self.partial_file_id (== Some(B)) != Some(file.id == A) → true → obsolete branch runs. skipped_test_names has no entry matching sibling (it was never sequenced), so "sibling 1" stays in unchecked_keys → self.obsolete += 1.

    Output: snapshots: … 1 obsolete + To remove obsolete snapshots, run bun test -u — for A's .only()-sibling snapshot, exactly the false positive partial_file_id was meant to prevent.

    There's also a green-run variant: if B uses test.only() on a test that completes before B's first toMatchSnapshot() (instead of a failure), the same overwrite happens on an exit-0 run, and neither the unhandled_errors_between_tests summary suppression (test_command.rs:~2861) nor a red exit code masks the false positive.

    Why existing code doesn't prevent it

    robobun's reply resolving the earlier backward-leak comment said "a later file setting Some(B) does not suppress file A since the ids do not match" — correct for the false-negative direction (B ≠ A ⇒ A's obsolete branch runs). But that same inequality is why A's own suppression is lost when A had set Some(A) first: the single slot can't remember that both A and B were partial. There is no reset between files, but a reset wouldn't help either — the write and the read both happen after B's overwrite.

    Impact

    Reporting-only false positive in the new obsolete counter. Narrow trigger: requires a specific multi-file ordering (A partial + opened its .snap; B's first completing sequence is fail/.only() and precedes B's first toMatchSnapshot()). In the failure variant the run is already red, so users are unlikely to blindly follow the -u hint; the .only()-in-both-files variant can hit on a green run but is a dev-loop state. No on-disk change without -u; following the hint under -u deletes the sibling entry, but that's pre-existing -u+.only() behaviour. Not merge-blocking.

    How to fix

    Either make it a small set — e.g. partial_file_ids: Vec<FileId> (or a bit on the per-file struct), pushed at Execution.rs:718 and checked with .contains(&file.id) at flush — or add an explicit write_snapshot_file() at the per-file boundary in the sequential loop so A is flushed before B's sequences start writing shared state. The Vec is the smaller change and mirrors how skipped_test_names already handles the same lazy-flush window.

Comment thread src/runtime/test_runner/snapshot.rs Outdated
Comment thread src/runtime/cli/test_command.rs Outdated
Comment thread src/runtime/cli/test_command.rs
…d errors

partial_file_ids: Vec<FileId> replaces the single Option slot so a later
file cannot overwrite an earlier file's own suppression across the lazy
flush window. mark_file_partial is also called where
unhandled_errors_between_tests is incremented (describe-callback throws
and errors between tests), so the suppression is per-file instead of the
previous run-wide summary gate. The hint-prefix arm gains a comment
noting the key-format ambiguity trade-off.
@robobun

robobun commented Jul 12, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed in a024016:

  • partial_file_id: Option<FileId> → partial_file_ids: Vec<FileId> with a mark_file_partial helper, so a later file can no longer overwrite an earlier file's own suppression across the lazy flush window.
  • mark_file_partial is also called from bun_test.rs where unhandled_errors_between_tests is incremented (covers throwing describe callbacks and stray rejections per-file). The run-wide summary gate from 4c2df2b is removed.
  • Added a code comment on the hint-prefix arm noting the key-format ambiguity trade-off.
  • --parallel snapshot counter aggregation is a pre-existing IPC gap (no snapshot counters have ever been sent worker→coordinator) and is out of scope here.

Comment thread src/runtime/test_runner/snapshot.rs
@robobun

robobun commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

Closing as part of a cleanup of stale pull requests. This PR has had no new commits since 2026-07-12, it conflicts with main, and its last CI run failed. This is not a judgment on the fix itself. The linked issue (#12114) stays open. If the problem still reproduces on a current build, reopen this PR after a rebase or open a new one against main.

@robobun robobun closed this Sep 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

On --update-snapshots, snapshot file remains empty in afterAll()

1 participant