Skip to content

bun:test: don't report beforeAll/afterAll as phantom '(unnamed)' tests in describe.skip/describe.todo - #35502

Open
robobun wants to merge 2 commits into
mainfrom
farm/83545e66/describe-skip-todo-phantom-hooks
Open

robobun wants to merge 2 commits into
mainfrom
farm/83545e66/describe-skip-todo-phantom-hooks

Conversation

@robobun

@robobun robobun commented Jul 25, 2026 •

Copy link
Copy Markdown
Collaborator

What

beforeAll/afterAll hooks inside a describe.skip or describe.todo block are reported as (unnamed) skipped/todo tests, inflating the summary and JUnit counts. beforeEach/afterEach are not affected.

import { describe, test, beforeAll, afterAll, beforeEach, afterEach } from "bun:test";
describe.skip("disabled suite", () => {
  beforeAll(() => {});
  afterAll(() => {});
  beforeEach(() => {});
  afterEach(() => {});
  test("the only real test", () => {});
});

Before:

(skip) disabled suite > (unnamed)
(skip) disabled suite > the only real test
(skip) disabled suite > (unnamed)

 3 skip
Ran 3 tests across 1 file.

JUnit: tests="3" skipped="3" with two <testcase name="(unnamed)"><skipped/></testcase> entries.
Jest 30: Tests: 1 skipped, 1 total.

Cause

describe.skip/describe.todo inherit their mode to every child, and ExecutionEntry::create drops the callback for Skip/Todo entries. Order::generate_all_order still scheduled each beforeAll/afterAll entry as its own ExecutionSequence with test_entry = None. At execution time the callback-less entry sets sequence.result to Skip/Todo, and on_sequence_completed reports any non-Pass sequence, falling back to first_entry (the hook, name = None) as the test to print.

beforeEach/afterEach are folded into each test's own sequence and only gathered when the test has a callback, so they never get a standalone sequence.

Fix

Skip hook entries whose callback is None in generate_all_order. A hook with no callback has nothing to run and is not a test to report. describe.todo with --todo is unaffected: run_todo keeps the callback, so the hook is still scheduled and runs.

After:

(skip) disabled suite > the only real test

 1 skip
Ran 1 test across 1 file.

JUnit: tests="1" skipped="1".

Related

#35497 removes always_use_hooks so an all-test.skip describe no longer runs its hooks, which also stops scheduling the describe.skip phantoms. It does not cover describe.todo (a todo test still marks its ancestors has_callback, so the nulled hooks are still scheduled). The two changes touch different code paths and compose.

Tests

Added three cases to test/cli/test/bun-test.test.ts asserting one reported test (no (unnamed), no hook output) for describe.skip and describe.todo, and that describe.todo with --todo still runs its beforeAll/afterAll.


no test proof · iteration 1 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/cli/test/bun-test.test.ts

A describe.skip/describe.todo block inherits its mode to every child,
which nulls the callback on its beforeAll/afterAll entries. generate_all_order
still scheduled each nulled hook as its own ExecutionSequence with
test_entry = None; on completion the sequence's result is Skip/Todo, so the
reporter prints it as an '(unnamed)' skipped/todo test and counts it in the
summary and JUnit output.

Skip hooks with no callback when building the schedule. describe.todo with
--todo is unaffected because run_todo keeps the callback.
@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 26 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9dcd0001-0d22-4d03-9661-1ad672b73961

📥 Commits

Reviewing files that changed from the base of the PR and between ae4b17d and 56a6abb.

📒 Files selected for processing (2)
  • src/runtime/test_runner/Order.rs
  • test/cli/test/bun-test.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Jul 25, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: diff is ready; waiting on a maintainer to rebuild or merge.

Reproduced with the snippet in the PR body: 3 skip / Ran 3 tests on canary, 1 skip / Ran 1 test with this change. bun bd test test/cli/test/bun-test.test.ts passes locally.

CI: test/cli/test/bun-test.test.ts (the changed test file) passed on every lane that ran across both builds.

  • build #79946: all lanes ran; only failures were three unrelated [flaky] tests (webview-chrome animation click, security-scanner-matrix TTY, next-pages dev-server connection refused).
  • build #80080 (retrigger): 151 jobs passed. Four build lanes timed out waiting for agents (darwin-aarch64, linux-x64-musl, freebsd-aarch64, windows-aarch64), windows-x64 test lane canceled on Azure agent-create failure, alpine-aarch64 test lanes expired. The one test-lane failure is test/cli/run/no-orphans.test.ts (perl daemon reaping timeout on darwin 14 x64), plus three [flaky] node-parallel/install tests. None touch the test-runner scheduling path this PR changes.

@robobun

robobun commented Jul 25, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 3:17 AM PT - Jul 25th, 2026

❌ @robobun, your commit 56a6abb has 2 failures in Build #80080 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 35502

That installs a local version of the PR into your bun-35502 executable, so you can run:

bun-35502 --bun

@github-actions

Copy link
Copy Markdown
Contributor

This PR may be a duplicate of:

  1. bun:test: skip beforeAll/afterAll for describe blocks with no non-skipped tests #35497 - Both fix beforeAll/afterAll hooks producing phantom "(unnamed)" tests in describe.skip/describe.todo blocks, both modify Order.rs in the test runner's hook scheduling logic

🤖 Generated with Claude Code

@robobun

robobun commented Jul 25, 2026

Copy link
Copy Markdown
Collaborator Author

Not a duplicate. #35497 gates hook scheduling on has_callback, and a test.todo child still marks its ancestors has_callback = true (it's not Skip/FilteredOut). So under #35497 a describe.todo block's beforeAll/afterAll are still scheduled with a nulled callback and still reported as (unnamed) todo tests. This PR skips callback-less hook entries in generate_all_order, which covers both describe.skip and describe.todo. The two changes touch different lines and compose cleanly (see the Related section in the PR body).

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — targeted fix that stops scheduling callback-less beforeAll/afterAll entries as standalone sequences.

What was reviewed:

  • Confirmed generate_all_order is only reached for before_all/after_all lists (describe scopes and the preload hook_scope); the preload root is Normal mode so its hooks always keep their callback and are unaffected.
  • Traced ExecutionEntry::create: callback is None only for Skip mode, or Todo without run_todo — matches the guard and the --todo test.
  • Checked AllOrderResult::set_failure_skip_to and previous_group_was_concurrent when every entry continues — both degrade to the same state as an empty hook list, so no ordering change.
Extended reasoning...

Overview

Single-statement change in src/runtime/test_runner/Order.rs: generate_all_order now continues past hook entries whose callback is None, so a beforeAll/afterAll inside describe.skip/describe.todo no longer gets its own ExecutionSequence (which the reporter was printing as an (unnamed) skipped/todo test). Three regression tests added to test/cli/test/bun-test.test.ts.

Security risks

None. This is test-runner scheduling metadata; no untrusted input, no allocation, no FFI surface touched.

Level of scrutiny

Low-to-medium. The guard is a pure early-continue on an Option::is_none() before any of the unsafe pointer manipulation in the loop body. I verified all four call sites of generate_all_order (two in generate_order_describe, two for the preload hook_scope in bun_test.rs) pass only hook lists, and cross-checked ExecutionEntry::create (bun_test.rs:1936-1945) to confirm callback is nulled exactly for Skip and for Todo-without-run_todo, so the --todo path correctly keeps scheduling. The preload root scope is created with Normal mode, so its hooks always retain a callback and never hit the new continue.

I also checked the two bits of loop-external state the skipped iteration would have touched: AllOrderResult with start == end iterates an empty slice in set_failure_skip_to, and leaving previous_group_was_concurrent untouched matches what already happens when before_all is empty — so a nulled hook now behaves identically to no hook, which is the intent.

Other factors

Tests cover both .skip and .todo, assert the negative (not.toContain("(unnamed)"), hooks not executed), the positive (exact 1 skip/1 todo count and Ran 1 test), and separately verify --todo still executes the hooks. They use the existing runTest helper and neighboring conventions. The PR description explains composition with #35497 (different code path). No prior reviewer comments to address.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — targeted one-line guard with clear tests.

Checked that generate_all_order is only reached for before_all/after_all slices (describe-scope and root preload); root/preload hooks can't inherit Skip/Todo mode so their callback is never None, and ExecutionEntry::create confirms callback is nulled exactly for Skip and Todo-without-run_todo. Verified AllOrderResult with start == end is a no-op in set_failure_skip_to, and that the --todo path keeps the callback so hooks still run (covered by the third test). The only side effect skipped by the early continue is previous_group_was_concurrent = false, which only matters for concurrent-group merging of a callback-less skipped test — benign since it never executes.

Extended reasoning...

Overview

Adds a single early continue in Order::generate_all_order (src/runtime/test_runner/Order.rs:45) that skips beforeAll/afterAll entries whose callback is None. These entries only get a None callback when their enclosing describe is .skip or .todo (without --todo), per ExecutionEntry::create. Previously each such hook was scheduled as its own ExecutionSequence with test_entry = None, which the reporter printed as an (unnamed) skipped/todo test and counted in the summary and JUnit output.

Three tests are added to test/cli/test/bun-test.test.ts covering describe.skip, describe.todo, and describe.todo with --todo (the positive case proving hooks still run when they should).

Security risks

None. This is test-runner scheduling logic with no external input handling, no auth/crypto, and no memory-safety changes — the new guard runs before any of the unsafe pointer manipulation in the loop body.

Level of scrutiny

Low-to-moderate. It's a 1-line behavioral change plus a comment in a hot-ish scheduling path, but the guard is narrow: callback.is_none() on a hook entry means there is literally nothing to execute and nothing to report. I traced all four call sites of generate_all_order (two in generate_order_describe, two for the root preload hook_scope in bun_test.rs) — preload hooks can't be inside a .skip/.todo so their callback is always Some, and the describe-scope calls are exactly the target of the fix.

Other factors

  • AllOrderResult with start == end (all entries skipped) makes set_failure_skip_to a no-op via an empty range — no off-by-one.
  • The --todo test guards against over-skipping: with run_todo, ExecutionEntry::create keeps the callback so the hook is still scheduled.
  • The only observable state not reset when all hook entries are skipped is previous_group_was_concurrent; this could theoretically let a concurrent skipped test extend a preceding concurrent group, but skipped tests don't execute and are reported identically either way.
  • Relationship to #35497 is documented in the PR body and the author's follow-up comment; the two touch different lines and compose.
  • Tests use console.error and runTest captures stderr (stdout: "ignore"), so the hook-ran assertions are wired correctly.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants