Skip to content

sql: union batch-insert keys across all rows; encode null as SQL NULL in sql.array - #35111

Open
robobun wants to merge 5 commits into
mainfrom
farm/6761c8bf/sql-helper-union-keys-array-null
Open

robobun wants to merge 5 commits into
mainfrom
farm/6761c8bf/sql-helper-union-keys-array-null

Conversation

@robobun

@robobun robobun commented Jul 22, 2026 •

Copy link
Copy Markdown
Collaborator

What

Two silent data-loss bugs in Bun.SQL helpers.

Batch insert drops keys not present on the first row

const rows = [
  { id: 1, a: "onlyA" },                     // no `b` key
  { id: 2, a: "hasB", b: "IMPORTANT-DATA" },
  { id: 3, b: "b-only" },
];
await sql`insert into h ${sql(rows)} returning id, a, b`;
// before: [{id:1,a:"onlyA",b:null},{id:2,a:"hasB",b:null},{id:3,a:null,b:null}]
// after:  [{id:1,a:"onlyA",b:null},{id:2,a:"hasB",b:"IMPORTANT-DATA"},{id:3,a:null,b:"b-only"}]

SQLHelper derived the column list from Object.keys(rows[0]) only, so a key that first appears on a later row was never emitted and its value was silently discarded. buildDefinedColumnsAndQuery already scans every row and null-fills missing values, so the fix is to union the keys across all rows when no explicit column list is given. Explicit columns (sql(rows, "a", "b")) and single-row / positional-array inputs are unchanged.

sql.array([..., null, ...]) stores the string "null"

sql.array(["a", null, "b"], "TEXT").serializedValues
// before: {"a","null","b"}   -> TEXT[] stores the four-character string "null"
// after:  {"a",null,"b"}     -> SQL NULL

sql.array([1, null, 2], "INTEGER").serializedValues
// before: {1,"null",2}       -> server rejects with 22P02 "invalid input syntax for type integer"
// after:  {1,null,2}

arrayValueSerializer had no branch for null; typeof null === "object" sent it through JSON.stringify and produced the quoted string "null". It now emits the unquoted null token (SQL NULL in Postgres array-literal syntax), matching the existing undefined handling.

Tests

test/js/sql/sql-helpers-validation.test.ts gains no-server coverage for both fixes (sqlite in-memory for the insert helper, serializedValues inspection for sql.array). test/js/sql/sql.test.ts gains live-Postgres round-trip tests under the docker guard.


[review] gate passed · iteration 2 · 4 files touched

fails on main (without fix)
ASAN without fix: 4 FAILED
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/sql/sql-helpers-validation.test.ts test/js/sql/sql.test.ts
bun test v1.4.0 (b210150fb)

test/js/sql/sql-helpers-validation.test.ts:
(pass) sqlite helper validation > null items in WHERE IN helper with a column are rejected [194.48ms]
(pass) sqlite helper validation > null and undefined items in INSERT helper are rejected [27.32ms]
(pass) sqlite helper validation > null and undefined items in UPDATE helper are rejected [21.02ms]
(pass) sqlite helper validation > empty update helper throws regardless of SET casing [31.67ms]
(pass) sqlite helper validation > empty update helper throws even alongside a literal assignment [26.39ms]
(pass) postgres helper validation > null items in WHERE IN helper with a column are rejected [46.28ms]
(pass) postgres helper validation > null and undefined items in INSERT helper are rejected [12.66ms]
(pass) postgres helper validation > null and undefined items in UPDATE helper are rejected [9.68ms]
(pass) postgres helper validation > empty update helper throws regardless of SET casing [16.01ms]
(pass)
... (truncated)

release without fix: 1 FAILED
bun test v1.4.0-canary.1 (3e0954da1)

test/js/sql/sql-helpers-validation.test.ts:
(pass) sqlite helper validation > null items in WHERE IN helper with a column are rejected [3.54ms]
(pass) sqlite helper validation > null and undefined items in INSERT helper are rejected [0.52ms]
(pass) sqlite helper validation > null and undefined items in UPDATE helper are rejected [0.30ms]
(pass) sqlite helper validation > empty update helper throws regardless of SET casing [0.41ms]
(pass) sqlite helper validation > empty update helper throws even alongside a literal assignment [0.30ms]
(pass) postgres helper validation > null items in WHERE IN helper with a column are rejected [1.02ms]
(pass) postgres helper validation > null and undefined items in INSERT helper are rejected [0.15ms]
(pass) postgres helper validation > null and undefined items in UPDATE helper are rejected [0.09ms]
(pass) postgres helper validation > empty update helper throws regardless of SET casing [0.24ms]
(pass) postgres helper validation > empty update helper throws even alongside a literal assignment [0.09ms]
(pass) mysql helper validation > null items in WHERE IN helper with a column are rejected [0.16ms]
... (truncated)
passes on PR (with fix)
ASAN with fix: all passed
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/sql/sql-helpers-validation.test.ts test/js/sql/sql.test.ts
bun test v1.4.0 (b210150fb)

test/js/sql/sql-helpers-validation.test.ts:
(pass) sqlite helper validation > null items in WHERE IN helper with a column are rejected [240.55ms]
(pass) sqlite helper validation > null and undefined items in INSERT helper are rejected [31.64ms]
(pass) sqlite helper validation > null and undefined items in UPDATE helper are rejected [27.11ms]
(pass) sqlite helper validation > empty update helper throws regardless of SET casing [44.50ms]
(pass) sqlite helper validation > empty update helper throws even alongside a literal assignment [27.47ms]
(pass) postgres helper validation > null items in WHERE IN helper with a column are rejected [48.25ms]
(pass) postgres helper validation > null and undefined items in INSERT helper are rejected [12.79ms]
(pass) postgres helper validation > null and undefined items in UPDATE helper are rejected [10.02ms]
(pass) postgres helper validation > empty update helper throws regardless of SET casing [17.11ms]
(pass
... (truncated)

release with fix: all passed
$ bun scripts/build.ts --profile=release
[configured] bun-profile → bun (stripped) in 786ms (unchanged)
ninja: Entering directory `/workspace/bun/build/release'
[1/21] gen JS modules (bundle-modules)
Preprocess modules (8542ms)
Bundle modules (36ms)
Postprocesss modules (161ms)
Bundle Functions (795ms)
Generate Code (27ms)

[9.57s] Bundled "src/js" for production
  2041 kb
  165 internal modules
  13 native modules
  90 internal functions across 19 files
[1/8] cargo bun_bin → libbun_rust.a (--target x86_64-unknown-linux-gnu)

  nightly-2026-07-20-x86_64-unknown-linux-gnu unchanged - rustc 1.99.0-nightly (9f36de775 2026-07-19)

�[1m�[92m   Compiling�[0m bun_core v0.0.0 (/workspace/bun/src/bun_core)
�[1m�[92m   Compiling�[0m bun_errno v0.0.0 (/workspace/bun/src/errno)
�[1m�[92m   Compiling�[0m bun_ptr v0.0.0 (/workspace/bun/src/ptr)
�[1m�[92m   Compiling�[0m bun_boringssl_sys v0.0.0 (/workspace/bun/src/boringssl_sys)
�[1m�[92m   Compiling�[0m bun_safety v0.0.0 (/workspace/bun/src/safety)
�[1m�[92m   Compiling�[0m bun_zlib_sys v0.0.0 (/workspace/bun/src/zlib_sys)
�[1m�[92m   Compiling�[0m bun_cares_sys v0.0.0 (/workspace/bun/src/cares_sys)
�[1m�[92m   Compiling�[0m
... (truncated)
diff hotspot
src/js/internal/sql/postgres.ts            |   6 ++
 src/js/internal/sql/shared.ts              |  89 ++++++++++++++++++-------
 test/js/sql/sql-helpers-validation.test.ts | 100 +++++++++++++++++++++++++++++
 test/js/sql/sql.test.ts                    |  37 +++++++++++
 4 files changed, 210 insertions(+), 22 deletions(-)

gate history · 4 passed · 0 rejected · iteration 2

evidence per changed file
file                                        reads  edits  tests
src/js/internal/sql/postgres.ts                 2      1      0
src/js/internal/sql/shared.ts                  11     10      0
test/js/sql/sql-helpers-validation.test.ts      4      4      0
test/js/sql/sql.test.ts                         2      3      0

… in sql.array

The sql(rows) batch-insert helper derived its column list from
Object.keys(rows[0]) only. A key that first appeared on a later row was
never emitted as a column, so its value was silently discarded and NULL
was stored for every row. The downstream buildDefinedColumnsAndQuery
already scans every row and null-fills missing values, so the column
list just needs to be the union of keys across all rows.

Separately, arrayValueSerializer had no branch for null: typeof null is
'object', so it fell through to JSON.stringify and produced the quoted
string '"null"'. A TEXT[] column stored the four-character string and
an INTEGER[] column rejected it with 22P02. Emit the unquoted null token
(SQL NULL in Postgres array-literal syntax), matching the existing
undefined handling.
@coderabbitai

coderabbitai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 2 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: aff468f3-53cc-4438-86f2-cb2be3c6824b

📥 Commits

Reviewing files that changed from the base of the PR and between 47597ab and b210150.

📒 Files selected for processing (4)
  • src/js/internal/sql/postgres.ts
  • src/js/internal/sql/shared.ts
  • test/js/sql/sql-helpers-validation.test.ts
  • test/js/sql/sql.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Jul 22, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: diff is green; all review threads resolved. Ready for review.

Verified with:

bun bd test test/js/sql/sql-helpers-validation.test.ts   # 31 pass / 0 fail

Fail-before (src/ reverted to main): batch insert helper unions keys across all rows and the three null in ... array serializes as SQL NULL cases fail as expected.

Follow-ups since the first push:

  • 198e234 scopes the key-union to the INSERT path only (via autoColumns), so WHERE IN / UPDATE behaviour is unchanged.
  • 3e0954d extracts validateHelperKey so keys contributed by later rows are validated the same as first-row keys.
  • b210150 lets the INSERT key-widening run when the first row is {} (the columnCount === 0 guard now defers to the post-widening check).

CI: no SQL test failed on any lane across builds #77719, #77758, #77766. Remaining failures are unrelated to this diff:

  • test/js/node/test/parallel/test-net-connect-memleak.js is failing consistently on main across all three builds (GC-timing assertion in node:net; this diff touches only src/js/internal/sql/; reported for triage).
  • The rest are [flaky] install/transpiler/bake tests that passed on retry on other lanes.

@robobun

robobun commented Jul 22, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 7:41 AM PT - Jul 22nd, 2026

❌ @robobun, your commit b210150 has 3 failures in Build #77766 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 35111

That installs a local version of the PR into your bun-35111 executable, so you can run:

bun-35111 --bun

@github-actions

Copy link
Copy Markdown
Contributor

Found 1 issue this PR may fix:

  1. Bug: Bun.SQL bulk-insert rows with nullable columns causes database to crash #28980 - PR fixes batch insert key discovery (unioning keys across all rows instead of only first row's keys) and null encoding in sql.array, directly addressing the bulk-insert with nullable columns scenario

If this is helpful, copy the block below into the PR description to auto-close this issue on merge.

Fixes #28980

🤖 Generated with Claude Code

@robobun

robobun commented Jul 22, 2026

Copy link
Copy Markdown
Collaborator Author

The suggested link to #28980 is a false positive. That issue is about the prepared-statement cache key including per-parameter null type suffixes (causing unbounded Parse growth when null patterns vary across batches). This PR does not touch the statement cache; it fixes column-list derivation in SQLHelper and null serialization in sql.array, which are independent code paths.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find any bugs. Both fixes look correct and the tests are solid, but the key-union change is a user-facing API semantics decision (heterogeneous rows now succeed with unioned columns rather than silently using row-0's keys) that's worth a maintainer sign-off.

What was reviewed:

  • arrayValueSerializer null branch — verified unquoted null is the correct Postgres array-literal token; MySQL's array() throws so no sibling site to patch.
  • SQLHelper union path — confirmed single-object ([obj]), positional-array rows, and explicit-column calls still hit the old Object.keys(value[0]) path; buildDefinedColumnsAndQuery already null-fills missing keys so downstream is unchanged.
  • Noted a minor edge: WHERE id IN ${sql([{id:1},{id:2,other:x}])} (no explicit column) previously used only id, now unions to two columns and throws "Cannot use WHERE IN helper with multiple columns". Unlikely in practice but a behavior change.
Extended reasoning...

Overview

Two independent fixes in Bun.SQL helpers:

  • src/js/internal/sql/postgres.ts: arrayValueSerializer gains an early value === null branch returning the unquoted token "null", matching the existing undefined handling. Without it, typeof null === "object" fell through to JSON.stringify and produced the quoted string "null".
  • src/js/internal/sql/shared.ts: SQLHelper's constructor now unions Object.keys across all rows when the caller passes a multi-row array of objects with no explicit column list, instead of deriving columns from rows[0] only.
  • Tests added in both the no-server helper-validation file (sqlite in-memory + serializedValues inspection) and the docker-guarded live-Postgres suite.

Security risks

None. Column names still flow through escapeIdentifier; array elements are still escaped via the existing arrayEscape. The new branch only affects the literal null value.

Level of scrutiny

Medium. These are small, targeted fixes to real data-loss bugs, but they change user-visible semantics of a public API (sql(rows) batch insert). The union-keys approach is one valid design; another would be to throw on heterogeneous rows (forcing the caller to pass explicit columns). A maintainer should confirm union-and-null-fill is the intended contract.

Other factors

  • Checked call sites: SQLHelper is constructed at three places in src/js/bun/sql.ts, all passing the user's array + rest-arg column names; the single-object path wraps in [obj] so length > 1 is false and it takes the unchanged branch.
  • The union loop skips null/non-object items rather than throwing; buildDefinedColumnsAndQuery still throws "Cannot use null or undefined as an item in INSERT helper" downstream, so the error surface is preserved.
  • MySQL adapter's array() throws "MySQL doesn't support arrays", so the null-serialization fix has no sibling to patch there.
  • One subtle knock-on: for WHERE ... IN ${sql(objs)} with no explicit column and heterogeneous objects, the derived column count can now exceed 1, which trips the existing "Cannot use WHERE IN helper with multiple columns" guard. Previously it silently used row-0's single key. This is edge-case usage (WHERE IN with objects normally takes an explicit column), and turning silent key-drop into an error is arguably safer, but it's a behavior change worth being aware of.
  • Test coverage is good: fail-before was verified, explicit-column path has a regression guard, and the live-Postgres round-trip asserts both TEXT[] and INTEGER[] null handling.

Move the union-of-keys logic out of the SQLHelper constructor and into
the INSERT branch of normalizeQuery, gated on a new autoColumns flag.
WHERE IN over heterogeneous objects with no explicit column now keeps
its existing first-row-keys behaviour instead of tripping the 'multiple
columns' guard. Adds a regression test for that case.
@robobun

robobun commented Jul 22, 2026

Copy link
Copy Markdown
Collaborator Author

Good catch on the WHERE IN edge case. Verified: on main, WHERE id IN ${sql([{id:1},{id:2,other:"x"}])} succeeds using the first row's id key; the original patch would have made it throw "Cannot use WHERE IN helper with multiple columns".

198e234 scopes the key-union to the INSERT path via a new autoColumns flag on SQLHelper. normalizeQuery now widens the column list only when building an INSERT with an auto-derived column list and more than one row, so WHERE IN and UPDATE are untouched. Added a regression test for the WHERE IN case.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline nit, I also checked: MySQL/SQLite have no sibling arrayValueSerializer needing the same null fix (MySQL's array() throws, SQLite has none); positional-array batch inserts (sql([[1,"a"],[2,"b"]])) are correctly excluded from key-widening by the !$isArray(items[0]) guard; and unionRowKeys' use of global Object.keys/Set matches the existing constructor code in the same file.

Extended reasoning...

The two fixes are correct and well-tested, and the WHERE IN regression from the first iteration was already scoped out with a covering test. Deferring rather than approving because this changes user-facing INSERT column-list derivation semantics in Bun.SQL — clearly the right change (the old behavior silently dropped data), but an API-surface behavior change worth a maintainer glance.

Comment thread src/js/internal/sql/shared.ts
Extract the SQLHelper constructor's key check into validateHelperKey and
run it on every key unionRowKeys adds from later rows, so a key like
"70000" is rejected regardless of which row it first appears on.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the validateHelperKey extraction addresses my prior note, and both fixes (key-union for INSERT, unquoted null in sql.array) are correct and well-covered. The remaining inline nit (empty first row still hits the pre-widening columnCount === 0 guard) is not a regression and fails loudly, so it's non-blocking.

Extended reasoning...

Overview

Two independent silent-data-loss fixes in Bun.SQL helpers: (1) src/js/internal/sql/shared.ts gains unionRowKeys + an autoColumns flag on SQLHelper so batch INSERT with an auto-derived column list unions keys across all rows instead of dropping keys absent from row 0; (2) src/js/internal/sql/postgres.ts adds a value === null branch to arrayValueSerializer so sql.array([..., null, ...]) emits SQL NULL instead of the quoted string "null". The constructor's key-validation loop was extracted into validateHelperKey and applied to widened keys per my prior review. Tests cover both fixes without a server (sqlite in-memory + serializedValues inspection) plus docker-guarded Postgres round-trips.

Security risks

None identified. Widened column names come from Object.keys() (own-enumerable strings only), pass through validateHelperKey, and are quoted by escapeIdentifier, so there is no injection vector. The null branch in arrayValueSerializer emits a fixed literal token. No auth, TLS, or credential paths touched.

Level of scrutiny

Medium — user-facing SQL query generation, but the changes are narrowly scoped and strictly correctness-improving. The key-union is gated by autoColumns && (items) && items.length > 1 && !(items[0]), so explicit-column, single-row, positional-array, WHERE IN, and UPDATE paths are all untouched (WHERE IN preservation has a dedicated regression test). The validateHelperKey refactor is a straight extraction of the prior inline logic with identical semantics.

Other factors

My previous review's nit (validation not applied to widened keys) was addressed in 3e0954d and the thread resolved. The new inline nit — [{}, {id:1}] still throws Cannot INSERT with no columns because the columnCount === 0 guard runs before widening — is not a regression (it threw identically before this PR) and fails loudly with a clear message, so it doesn't block. Fail-before/pass-after evidence is in the PR body; the no-server tests in sql-helpers-validation.test.ts run on all three adapters and don't depend on the docker guard.

Comment thread src/js/internal/sql/shared.ts
The shared columnCount === 0 guard fired before unionRowKeys, so
sql([{}, {id:1}]) still threw 'Cannot INSERT with no columns'. Skip the
guard for INSERT (it already has its own definedColumnCount === 0 check
after widening) so an empty first row is covered by the same widening as
any other missing-key row. Adds test coverage for the empty-first-row
and all-empty boundary cases.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both earlier inline comments are addressed and I didn't find anything further, but this changes the user-facing semantics of sql(rows) for heterogeneous row shapes (union-of-keys instead of first-row-only) and relaxes the columnCount === 0 guard for INSERT — worth a maintainer sign-off on the API choice.

What was reviewed:

  • validateHelperKey extraction is byte-equivalent to the old inline loop; widened keys now go through it.
  • Widening is scoped to autoColumns && multi-row && !array-row, so explicit columns / single-row / positional-array / WHERE IN / UPDATE paths are unchanged.
  • INSERT-with-zero-columns now falls through to the post-widening definedColumnCount === 0 throw (message changes for the single-{}-row case; still fails loudly).
  • arrayValueSerializer null branch precedes the $isArray check so nested-array recursion still handles inner nulls; undefined path unchanged.
Extended reasoning...

Overview

Two fixes in src/js/internal/sql/: (1) SQLHelper now records whether its column list was auto-derived (autoColumns), and the INSERT branch of normalizeQuery widens that list to the union of own keys across all rows via a new unionRowKeys helper, so keys first appearing on a later row are no longer silently dropped. The constructor's key-validation loop is extracted to validateHelperKey and applied to widened keys as well. (2) arrayValueSerializer in postgres.ts gains a value === null branch so sql.array emits the unquoted null token instead of the JSON-stringified "null". Tests in sql-helpers-validation.test.ts (no-server sqlite/serialization) and sql.test.ts (docker-gated live Postgres) cover both.

Security risks

None identified. Column names come from Object.keys() (own-enumerable strings only, no prototype walk) and are routed through escapeIdentifier; validateHelperKey gates the widened set the same as the seed set. The sql.array change only narrows what reaches JSON.stringify.

Level of scrutiny

Medium-high: this is shared query-normalization code on the hot path for all three SQL adapters, and it changes user-visible behavior of a public helper. The mechanics are small and well-tested, but the design choice — silently unioning keys vs. requiring explicit columns for heterogeneous batches (or matching whatever postgres.js does) — is an API decision a maintainer should confirm.

Other factors

The PR went through three follow-up commits addressing review feedback (scoping widening to INSERT only, validating widened keys, and letting an empty first row fall through to widening). Test coverage includes the negative contracts (explicit columns still ignore extras, WHERE IN unchanged, all-rows-empty still throws). One incidental behavior change: INSERT ... ${sql([{}])} now throws "Insert needs to have at least one column with a defined value" instead of "Cannot INSERT with no columns" — a message change only. I did not find any correctness issues in the current revision.

@robobun

robobun commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

#41301 overlaps with the postgres.ts hunk of this PR. It also binds a null element in sql.array as SQL NULL. One difference: in an explicit JSON or JSONB array, #41301 keeps null as the JSON value null, as on main.

If #41301 merges first, this PR still needs its batch-insert change in shared.ts.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants