From a2898b15927ab99483968b9e1cdf9a9a2cd4935a Mon Sep 17 00:00:00 2001 From: Joel Natividad <1980690+jqnatividad@users.noreply.github.com> Date: Sat, 19 Sep 2026 13:29:08 -0400 Subject: [PATCH 1/4] test: cover the Polars SQL changes since py-1.44.0 (28 sqlp tests) 50c2cf18f bumped Polars to rev 9d5804d. 20 commits touched crates/polars-sql between upstream py-1.44.0 (9ae4e57f5) and that pin, and none of them had test coverage here: the previous bump (6d33fb41f, to d84c1d4) added zero new test functions to tests/test_sqlp.rs, only adjusting existing ones for join-order fallout. So the whole py-1.44.0..9d5804d window was uncovered, not just the last few days of it. Adds 28 tests and 6 shared fixtures (sqlp suite 106 -> 134). Every expectation was captured by running the freshly rebuilt binary and pinning its actual output -- upstream's own tests drive SQLContext directly, so their SQL syntax transfers but their expected values do not (qsv layers CSV type inference, null rendering and float formatting on top). New syntax and functions: - GROUPING SETS / ROLLUP / CUBE, GROUPING(), GROUPING_ID() (#29278) - the bare date-part functions YEAR/QUARTER/MONTH/WEEK/DAY/ DAYOFMONTH/DAYOFWEEK/DAYOFYEAR/HOUR/MINUTE/SECOND (#29269) - APPROX_QUANTILE, 2-4 args (#29288) - typed DATE '...' / TIMESTAMP '...' literals (#29007) - date +/- integer arithmetic; Decimal non-equi joins (#29156) - parenthesized JOIN ... ON, and aliases inside those parens (#28967, #29158) - OVER on multi-argument aggregates (#29160) - NULLS FIRST/LAST inside a window's own ORDER BY (#29159) - EXISTS/NOT EXISTS, correlated scalar subqueries, CTE shadowing, case-insensitive relation names, ORDER BY over an unselected aggregate (#29006, #29010) Behavior changes pinned: - CAST( AS DATE/TIME/TIMESTAMP) parses instead of casting, and is strict where TRY_CAST is not. Two distinct failures: some rows parse -> "conversion from `str` to `date` failed"; none parse -> "could not find an appropriate format to parse dates". (#28062, #28986) - Postgres scope strictness: an aliased relation's original name is out of scope, so SELECT t1.a FROM t1 AS f now errors. (#28937) - a scalar subquery's aggregate binds to its own relation. (#28939) - unaliased constants with GROUP BY are named literal, literal:1, literal:2 rather than colliding. (#29367) Cargo.toml enables polars' `approx_quantile` feature. Upstream cfg-gates the function and ships the feature only inside its docs-selection/full umbrellas, neither of which qsv uses, so sqlp rejected APPROX_QUANTILE outright before this. Verified empirically, not inferred: "unsupported function 'approx_quantile'". qsvmcp and qsvdp share the polars dep entry and both still build; qsvlite has no polars. Two traps worth recording, both of which produced a test that asserted nothing until it was fixed: - A projection of ONLY a nulled column writes the row as an empty line, which the test CSV reader then drops -- so a TRY_CAST assertion could not tell a nulled row from a missing one. The cast fixtures carry an `id` column for exactly this reason. - sqlp_join_non_equi_decimal originally joined on whole numbers, where an integer comparison gives the same answer and the test proved nothing about Decimal. It now uses fractional values where 1.05 < 1.10 is true but any integer truncation makes it false. sqlp has no --maintain-order and its join output order has been nondeterministic since 6d33fb41f, so every join test orders by a unique key rather than relying on an incidental order; each was also confirmed byte-identical over 12 runs. APPROX_QUANTILE is sketch-based, so its values are pinned exactly only after confirming byte-identical output over 20 runs -- a tolerance band against QUANTILE_CONT would have hidden numeric drift. Verified: full suite 3,958 passed / 0 failed; cargo t sqlp 5x green; 13 mutation tests on the least-obvious assertions, all 13 caught; double-run-check clean; no new clippy warnings. Co-Authored-By: Claude Opus 5 (1M context) --- CHANGELOG.md | 12 + Cargo.lock | 1 + Cargo.toml | 5 + tests/test_sqlp.rs | 1254 ++++++++++++++++++++++++++++++++++++++++++++ 4 files changed, 1272 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9e49cc0fbd..46f885f772 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,10 +7,22 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] ### Added +- **`sqlp`: new SQL surface from the Polars bump to `9d5804d`** (upstream `py-1.44.0` -> the pinned rev). None of this needed qsv changes beyond the bump, but none of it was covered by tests either, so `tests/test_sqlp.rs` now exercises each one: + - **`GROUP BY GROUPING SETS (...)` / `ROLLUP(...)` / `CUBE(...)`**, plus **`GROUPING(k, ...)`** and **`GROUPING_ID(k, ...)`** to tell a subtotal row's NULL marker apart from a NULL in the data ([#29278](https://github.com/pola-rs/polars/pull/29278)). Both pack one bit per argument, MSB first, so argument order matters. Pair them with `--wnull-value` - a subtotal's NULL keys are otherwise written as empty fields. + - **the bare date-part functions `YEAR`, `QUARTER`, `MONTH`, `WEEK`, `DAY`/`DAYOFMONTH`, `DAYOFWEEK`, `DAYOFYEAR`, `HOUR`, `MINUTE`, `SECOND`** ([#29269](https://github.com/pola-rs/polars/pull/29269)). Both are ISO-based, which is the part worth knowing: `WEEK('2024-12-31')` is **1** (ISO week 1 of 2025), not 53, and `DAYOFWEEK` counts Monday as 1. + - **`DATE '...'` / `TIMESTAMP '...'` typed literals** ([#29007](https://github.com/pola-rs/polars/pull/29007)), and **date +/- integer arithmetic** shifting by whole days with the integer on either side, from a literal or a column, plus **Decimal comparisons in non-equi joins** ([#29156](https://github.com/pola-rs/polars/pull/29156)). + - **parenthesized `JOIN ... ON (...)` constraints** ([#28967](https://github.com/pola-rs/polars/pull/28967)); relation aliases declared *inside* those parens now resolve in the outer SELECT, since bare parens around a join do not open a new scope ([#29158](https://github.com/pola-rs/polars/pull/29158)). + - **`OVER` on multi-argument aggregates** - `CORR`, `COVAR_POP`, `COVAR_SAMP`, `QUANTILE_CONT`, `QUANTILE_DISC`, `STRING_AGG` ([#29160](https://github.com/pola-rs/polars/pull/29160)) - and **`NULLS FIRST`/`NULLS LAST` inside a window's own `ORDER BY`** ([#29159](https://github.com/pola-rs/polars/pull/29159)). + - **`EXISTS`/`NOT EXISTS` and correlated scalar subqueries**, CTEs that shadow a registered table, case-insensitive relation names, and `ORDER BY` over an aggregate that is not in the SELECT list ([#29006](https://github.com/pola-rs/polars/pull/29006), [#29010](https://github.com/pola-rs/polars/pull/29010)). +- `sqlp`: **`APPROX_QUANTILE(col, q [, allowed_rank_error [, method]])`** - an approximate quantile aggregate, new to the Polars SQL frontend in [pola-rs/polars#29288](https://github.com/pola-rs/polars/pull/29288). Upstream cfg-gates it behind a `approx_quantile` feature that ships only inside their `docs-selection`/`full` umbrellas, so qsv now enables it explicitly - without that, `sqlp` rejected the function as `unsupported function 'approx_quantile'`. `allowed_rank_error` defaults to `0.001` and `method` selects the sketch (e.g. `'kll'`); anything outside 2-4 arguments is a syntax error rather than a silent fallback. Being sketch-based it trades exactness for speed on large inputs - use `QUANTILE_CONT`/`QUANTILE_DISC` when you need the exact value. - **gallery: a real per-capita county rate map where the rate does NOT invert the ranking.** The gallery's two existing `--denominator census` figures both show a raw-count choropleth whose ranking *inverts* once divided by population - `district_requests` on six hand-drawn rectangles, `tristate_county_incidents` on 210 real counties with counts constructed to be anti-correlated with population. Both teach the lesson well and both are synthetic in exactly the place that matters, so a reader could fairly conclude that per-capita always flips a map. `wpa_211_requests` is the counter-example, on data that is real all the way down: **279,464 2-1-1 helpline requests** across the **25 Western Pennsylvania counties** served by the United Way of Southwestern Pennsylvania - the upstream resource *entire*, 2023-09-22 to 2025-08-27, no window and no sample (via the [WPRDC](https://data.wprdc.org/dataset/211-requests)). The count panel is Allegheny County and almost nothing else - **172,035 requests, 61.6% of the file** - and the rate panel beside it leaves Allegheny on top at **138.9 per 1,000 residents**. What moves is the middle: **Venango climbs 11th to 3rd**, **Cambria 5th to 2nd**, **Butler falls 8th to 13th**, and the two top-eights share only 6 of 8 members. Per-capita is a *different question*, not a trick that flips a map. The figure also documents the caveat the other two cannot: a rate map of a *helpline* measures service reach as much as need, and the map says so itself: its darkest county is **McKean, at 2 requests against 39,904 residents** - 0.05 per 1,000, where the next-lowest county (Elk) is 12.3, some 245 times higher. Nobody believes McKean has no hardship; it sits at the edge of *this* call center's intake, so the figure is measuring who dials 2-1-1, not who needs it. A per-capita map inherits whatever its numerator was actually counting. Nothing is stored in the CSV but the requests themselves: `--geojson auto` fetches the county boundaries **and** canonicalizes the county *names* to Census GEOIDs, and `--denominator census@2024` fetches ACS 5-year total population (`B01003`), so there is no FIPS column and no committed boundary file. The upstream feed's `zip_code` is dropped rather than mis-tagged (a second geo concept would add a competing choropleth candidate), and the dataset's older 2020-2023 resource is deliberately not concatenated: 5.4% of its rows carry a pipe-delimited `Allegheny County|Westmoreland County` multi-county value, and every way of resolving those changes the map. Built by the new `examples/viz/gen_wpa_211_requests.py`, which asserts its own row count, county count and date span so a silently revised upstream feed fails instead of quietly re-cutting the figure the caption describes. - `fetch` & `fetchpost`: **`--default-encoding `** - the fallback character encoding used to decode a response body when the server sends **no** `charset` parameter in its `Content-Type` header. Takes [WHATWG encoding labels](https://encoding.spec.whatwg.org/#names-and-labels) (`utf-8`, `windows-1252`, `iso-8859-1`, `shift_jis`, …) and defaults to `utf-8`, so behavior is unchanged unless you ask for it. A `charset` the server *does* send always wins; this only fills the gap for legacy APIs that serve latin-1/windows-1252 with a bare `Content-Type`. Unknown labels are rejected up front rather than silently falling back to UTF-8. The value participates in the cross-session disk/Redis cache key, so re-running the same URL under a different encoding re-decodes instead of serving the previous run's text. **Adding the encoding to that key invalidates existing `fetch`/`fetchpost` disk and Redis cache entries**, including for users who never pass the flag - a warm cache is re-fetched once after upgrading, and the orphaned entries age out on their normal TTL. ### Changed +- **`sqlp`: three Polars SQL behavior changes that can break an existing query.** All three arrive with the bump to `9d5804d` and are now pinned by tests: + - **`CAST( AS DATE/TIME/TIMESTAMP)` parses rather than casts.** The string->temporal cast kernel was removed upstream ([#28062](https://github.com/pola-rs/polars/pull/28062)) and SQL casts were re-routed through format inference ([#28986](https://github.com/pola-rs/polars/pull/28986)), so the spelling still works - but `CAST` is now **strict** and fails the whole query on a value it cannot parse, where `TRY_CAST` nulls just that row. The failure has two shapes: if *some* rows parse you get ``conversion from `str` to `date` failed``, naming the offending value; if *none* parse you get `could not find an appropriate format to parse dates`. + - **SQL scope rules follow Postgres** ([#28937](https://github.com/pola-rs/polars/pull/28937)). Once a relation is aliased, its original name is out of scope: `SELECT t1.a FROM t1 AS f` now errors with `no table or struct column named 't1' found`. An unqualified `ORDER BY` key still resolves, and an outer alias is still visible to a correlated subquery. + - **a scalar subquery's aggregate binds to its own relation** ([#28939](https://github.com/pola-rs/polars/pull/28939)), so aggregating an outer-relation column inside the subquery is an error instead of silently resolving; and **unaliased constants in a `SELECT` with `GROUP BY`** are projected once per group as `literal`, `literal:1`, `literal:2`, ... rather than colliding on one column ([#29367](https://github.com/pola-rs/polars/pull/29367)). - **`reqwest`'s `charset` and `system-proxy` features are now enabled explicitly.** Both were already active for the full `qsv` binary, but only by accident of Cargo feature unification: `cpc` (via `apply`), `plotly_static` (via `viz_static`) and `polars-io` all depend on `reqwest` *without* `default-features = false`, and reqwest's defaults include them. Listing them on qsv's own dependency pins that behavior, and turns it on for the leaner binaries that nothing else was pulling it into - **`qsvlite`, `qsvdp` and `qsvmcp` now**: - decode HTTP response bodies according to the `charset` parameter of the response's `Content-Type` (with BOM sniffing and stripping) instead of a blind lossy-UTF-8 conversion, and - honor **OS-level proxy configuration** (macOS SystemConfiguration, Windows registry) in addition to the `HTTP_PROXY`/`HTTPS_PROXY` environment variables. diff --git a/Cargo.lock b/Cargo.lock index 5732cee068..9fc3277536 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -6261,6 +6261,7 @@ dependencies = [ "rayon", "recursive", "regex", + "serde", "version_check", ] diff --git a/Cargo.toml b/Cargo.toml index 5ab65cfa1d..d8223494b4 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -317,6 +317,11 @@ rust-i18n = { version = "4.2.1", optional = true } pragmastat = "14.0" polars-readstat-rs = { version = "0.23", optional = true } polars = { version = "0.55", features = [ + # "approx_quantile" powers the APPROX_QUANTILE SQL aggregate in sqlp. It is + # NOT in any umbrella qsv enables (upstream keeps it in docs-selection/full + # only), so without it polars-sql cfg-gates the function out entirely and + # sqlp rejects it as "unsupported function 'approx_quantile'". + "approx_quantile", "asof_join", "avro", # "avx512" is x86_64-only — enabled via the diff --git a/tests/test_sqlp.rs b/tests/test_sqlp.rs index 7f4abb0299..8fef187aeb 100644 --- a/tests/test_sqlp.rs +++ b/tests/test_sqlp.rs @@ -4645,3 +4645,1257 @@ fn sqlp_invalid_format_errors() { wrk.assert_err(&mut cmd); } + +// --------------------------------------------------------------------------- +// Polars SQL surface added between py-1.44.0 and rev 9d5804d (the bump in +// 50c2cf18f). One test per upstream behavior change; the upstream PR is cited +// so the next bump can diff against it. +// --------------------------------------------------------------------------- + +/// Fixture shared by the GROUP BY / aggregate tests below. +fn grouping_fixture(wrk: &Workdir) { + wrk.create( + "groups.csv", + vec![ + svec!["category", "class", "value"], + svec!["a", "x", "1"], + svec!["a", "x", "2"], + svec!["a", "y", "3"], + svec!["b", "x", "4"], + svec!["b", "y", "5"], + svec!["b", "y", "6"], + ], + ); +} + +#[test] +fn sqlp_grouping_sets() { + // pola-rs/polars#29278: GROUP BY GROUPING SETS. The subtotal rows carry a + // NULL in the columns they do not group by, so --wnull-value makes the + // difference between "subtotal" and "a group whose key is empty" visible. + let wrk = Workdir::new("sqlp_grouping_sets"); + grouping_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("groups.csv") + .arg( + "SELECT category, class, SUM(value) AS total, COUNT(*) AS n FROM groups GROUP BY \ + GROUPING SETS ((category, class), (category), (class), ()) ORDER BY category NULLS \ + LAST, class NULLS LAST", + ) + .args(["--wnull-value", "NULL"]); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["category", "class", "total", "n"], + svec!["a", "x", "3", "2"], + svec!["a", "y", "3", "1"], + svec!["a", "NULL", "6", "3"], + svec!["b", "x", "4", "1"], + svec!["b", "y", "11", "2"], + svec!["b", "NULL", "15", "3"], + svec!["NULL", "x", "7", "3"], + svec!["NULL", "y", "14", "3"], + svec!["NULL", "NULL", "21", "6"], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_group_by_rollup() { + // pola-rs/polars#29278: ROLLUP(a, b) == GROUPING SETS ((a,b), (a), ()), + // i.e. it does NOT include the (class) subtotal that CUBE adds below. + let wrk = Workdir::new("sqlp_group_by_rollup"); + grouping_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("groups.csv") + .arg( + "SELECT category, class, SUM(value) AS total FROM groups GROUP BY ROLLUP(category, \ + class) ORDER BY category NULLS LAST, class NULLS LAST", + ) + .args(["--wnull-value", "NULL"]); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["category", "class", "total"], + svec!["a", "x", "3"], + svec!["a", "y", "3"], + svec!["a", "NULL", "6"], + svec!["b", "x", "4"], + svec!["b", "y", "11"], + svec!["b", "NULL", "15"], + svec!["NULL", "NULL", "21"], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_group_by_cube() { + // pola-rs/polars#29278: CUBE(a, b) adds the (class) subtotals that ROLLUP + // omits -- the two NULL-category rows below are the whole point. + let wrk = Workdir::new("sqlp_group_by_cube"); + grouping_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("groups.csv") + .arg( + "SELECT category, class, SUM(value) AS total FROM groups GROUP BY CUBE(category, \ + class) ORDER BY category NULLS LAST, class NULLS LAST", + ) + .args(["--wnull-value", "NULL"]); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["category", "class", "total"], + svec!["a", "x", "3"], + svec!["a", "y", "3"], + svec!["a", "NULL", "6"], + svec!["b", "x", "4"], + svec!["b", "y", "11"], + svec!["b", "NULL", "15"], + svec!["NULL", "x", "7"], + svec!["NULL", "y", "14"], + svec!["NULL", "NULL", "21"], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_grouping_function() { + // pola-rs/polars#29278: GROUPING(col) is 1 when col is rolled up in that + // row, else 0; with several arguments the bits are packed MSB-first, so + // GROUPING(category, class) == 3 on the grand total but 1 when only class + // is rolled up. + let wrk = Workdir::new("sqlp_grouping_function"); + grouping_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("groups.csv") + .arg( + "SELECT category, class, SUM(value) AS total, GROUPING(category) AS gc, \ + GROUPING(class) AS gk, GROUPING(category, class) AS gb FROM groups GROUP BY \ + ROLLUP(category, class) ORDER BY category NULLS LAST, class NULLS LAST", + ) + .args(["--wnull-value", "NULL"]); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["category", "class", "total", "gc", "gk", "gb"], + svec!["a", "x", "3", "0", "0", "0"], + svec!["a", "y", "3", "0", "0", "0"], + svec!["a", "NULL", "6", "0", "1", "1"], + svec!["b", "x", "4", "0", "0", "0"], + svec!["b", "y", "11", "0", "0", "0"], + svec!["b", "NULL", "15", "0", "1", "1"], + svec!["NULL", "NULL", "21", "1", "1", "3"], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_grouping_function_select_alias() { + // pola-rs/polars#29278: a SELECT alias over an expression is accepted both + // as the ROLLUP key and as the GROUPING() argument. + let wrk = Workdir::new("sqlp_grouping_function_select_alias"); + grouping_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("groups.csv") + .arg( + "SELECT UPPER(category) AS cat, SUM(value) AS total, GROUPING(cat) AS g FROM groups \ + GROUP BY ROLLUP(cat) ORDER BY cat NULLS LAST", + ) + .args(["--wnull-value", "NULL"]); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["cat", "total", "g"], + svec!["A", "6", "0"], + svec!["B", "15", "0"], + svec!["NULL", "21", "1"], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_date_integer_arithmetic() { + // pola-rs/polars#29156: a Date +/- an integer shifts by whole days, with the + // integer on either side and from either a literal or a column. The fixture + // straddles a leap day (2020-02-28 + 2 == 2020-03-01) and the same date in a + // non-leap year (2021-02-28 + 2 == 2021-03-02). + let wrk = Workdir::new("sqlp_date_integer_arithmetic"); + wrk.create( + "dates.csv", + vec![ + svec!["dt", "n"], + svec!["2020-01-01", "5"], + svec!["2020-02-28", "2"], + svec!["2021-02-28", "2"], + ], + ); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("dates.csv") + .arg( + "SELECT dt, n, dt + 5 AS plus_lit, 5 + dt AS lit_plus, dt - 5 AS minus_lit, dt + n AS \ + plus_col, dt - n AS minus_col FROM dates", + ) + .arg("--try-parsedates"); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec![ + "dt", + "n", + "plus_lit", + "lit_plus", + "minus_lit", + "plus_col", + "minus_col" + ], + svec![ + "2020-01-01", + "5", + "2020-01-06", + "2020-01-06", + "2019-12-27", + "2020-01-06", + "2019-12-27" + ], + svec![ + "2020-02-28", + "2", + "2020-03-04", + "2020-03-04", + "2020-02-23", + "2020-03-01", + "2020-02-26" + ], + svec![ + "2021-02-28", + "2", + "2021-03-05", + "2021-03-05", + "2021-02-23", + "2021-03-02", + "2021-02-26" + ], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_typed_temporal_literals() { + // pola-rs/polars#29007: `DATE '...'` / `TIMESTAMP '...'` are parsed as typed + // literals rather than a string that later gets cast, so they compare + // directly against a Date column. + let wrk = Workdir::new("sqlp_typed_temporal_literals"); + wrk.create( + "dates.csv", + vec![ + svec!["dt"], + svec!["2020-01-01"], + svec!["2020-02-28"], + svec!["2021-02-28"], + ], + ); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("dates.csv") + .arg("SELECT dt FROM dates WHERE dt > DATE '2020-01-15' ORDER BY dt") + .arg("--try-parsedates"); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![svec!["dt"], svec!["2020-02-28"], svec!["2021-02-28"]] + ); + + let mut between_cmd = wrk.command("sqlp"); + between_cmd + .arg("dates.csv") + .arg( + "SELECT dt FROM dates WHERE dt BETWEEN DATE '2020-01-01' AND DATE '2020-06-01' ORDER \ + BY dt", + ) + .arg("--try-parsedates"); + let got_between: Vec> = wrk.read_stdout_on_success(&mut between_cmd); + assert_eq!( + got_between, + vec![svec!["dt"], svec!["2020-01-01"], svec!["2020-02-28"]] + ); + + // As projected values: DATE keeps day resolution, TIMESTAMP is a Datetime. + let mut literal_cmd = wrk.command("sqlp"); + literal_cmd.arg("dates.csv").arg( + "SELECT DATE '2020-02-29' AS d, TIMESTAMP '2020-01-01 08:00:00' AS ts FROM dates LIMIT 1", + ); + let got_literal: Vec> = wrk.read_stdout_on_success(&mut literal_cmd); + assert_eq!( + got_literal, + vec![ + svec!["d", "ts"], + svec!["2020-02-29", "2020-01-01T08:00:00.000000"] + ] + ); +} + +#[test] +fn sqlp_cast_string_to_temporal() { + // pola-rs/polars#28062 removed the string->temporal *cast*, and #28986 then + // made `CAST( AS DATE/TIME/TIMESTAMP)` in SQL *parse* the string + // instead. So the SQL spelling keeps working even though the underlying + // cast is gone -- this test pins that, for a column operand, the `::` form, + // and a bare string literal. + let wrk = Workdir::new("sqlp_cast_string_to_temporal"); + wrk.create( + "strs.csv", + vec![ + svec!["d", "ts", "t"], + svec!["2000-02-01", "2000-02-01 12:30:00", "12:30:00"], + ], + ); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("strs.csv").arg( + "SELECT CAST(d AS DATE) AS d1, CAST(ts AS TIMESTAMP) AS ts1, CAST(ts AS DATETIME) AS ts2, \ + CAST(t AS TIME) AS t1 FROM strs", + ); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["d1", "ts1", "ts2", "t1"], + svec![ + "2000-02-01", + "2000-02-01T12:30:00.000000", + "2000-02-01T12:30:00.000000", + "12:30:00.000000000" + ] + ] + ); + + let mut colon_cmd = wrk.command("sqlp"); + colon_cmd + .arg("strs.csv") + .arg("SELECT d::date AS d1, ts::timestamp AS ts1, t::time AS t1 FROM strs"); + let got_colon: Vec> = wrk.read_stdout_on_success(&mut colon_cmd); + assert_eq!( + got_colon, + vec![ + svec!["d1", "ts1", "t1"], + svec![ + "2000-02-01", + "2000-02-01T12:30:00.000000", + "12:30:00.000000000" + ] + ] + ); + + let mut literal_cmd = wrk.command("sqlp"); + literal_cmd + .arg("strs.csv") + .arg("SELECT CAST('2000-02-01' AS DATE) AS d1, CAST('12:30:00' AS TIME) AS t1 FROM strs"); + let got_literal: Vec> = wrk.read_stdout_on_success(&mut literal_cmd); + assert_eq!( + got_literal, + vec![svec!["d1", "t1"], svec!["2000-02-01", "12:30:00.000000000"]] + ); +} + +#[test] +fn sqlp_cast_strict_vs_try_temporal() { + // pola-rs/polars#28986: CAST is strict and TRY_CAST is not. On a column + // holding one unparseable value, CAST fails the whole query while TRY_CAST + // nulls just that row. `sqlp_try_cast` covers only the TRY_CAST half. + let wrk = Workdir::new("sqlp_cast_strict_vs_try_temporal"); + // The `id` column is load-bearing: a projection of the nulled column ALONE + // writes the failed row as an empty line, which the CSV reader then drops, + // so the TRY_CAST assertion below could not tell a nulled row from a + // missing one. + wrk.create( + "badstrs.csv", + vec![ + svec!["id", "s"], + svec!["1", "2000-02-01"], + svec!["2", "not-a-date"], + ], + ); + + let mut strict_cmd = wrk.command("sqlp"); + strict_cmd + .arg("badstrs.csv") + .arg("SELECT id, CAST(s AS DATE) AS d FROM badstrs ORDER BY id"); + let stderr = wrk.stderr_on_error(&mut strict_cmd); + assert!( + stderr.contains("conversion from `str` to `date` failed"), + "unexpected stderr: {stderr}" + ); + assert!( + stderr.contains("not-a-date"), + "error should name the offending value: {stderr}" + ); + + let mut try_cmd = wrk.command("sqlp"); + try_cmd + .arg("badstrs.csv") + .arg("SELECT id, TRY_CAST(s AS DATE) AS d FROM badstrs ORDER BY id"); + let got: Vec> = wrk.read_stdout_on_success(&mut try_cmd); + assert_eq!( + got, + vec![svec!["id", "d"], svec!["1", "2000-02-01"], svec!["2", ""]] + ); +} + +#[test] +fn sqlp_approx_quantile() { + // pola-rs/polars#29288 added APPROX_QUANTILE to the SQL frontend. It is + // cfg-gated on polars' `approx_quantile` feature, which qsv enables + // explicitly in Cargo.toml -- upstream ships it only inside the + // docs-selection/full umbrellas, so without that line sqlp rejects the + // function outright. + // + // Signatures: (col, q), (col, q, allowed_rank_error), (col, q, error, method). + // The sketch is deterministic for a given input, so these are exact + // expectations, not a tolerance band (verified byte-identical over 20 runs). + let wrk = Workdir::new("sqlp_approx_quantile"); + grouping_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("groups.csv").arg( + "SELECT APPROX_QUANTILE(value, 0.5) AS a2, APPROX_QUANTILE(value, 0.5, 0.01) AS a3, \ + APPROX_QUANTILE(value, 0.5, 0.01, 'kll') AS a4, APPROX_QUANTILE(value, 0.25) AS q25, \ + APPROX_QUANTILE(value, 0.75) AS q75 FROM groups", + ); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["a2", "a3", "a4", "q25", "q75"], + svec!["4", "4", "4", "2", "5"] + ] + ); + + // It is a real aggregate, so it works per group. + let mut grouped_cmd = wrk.command("sqlp"); + grouped_cmd.arg("groups.csv").arg( + "SELECT category, APPROX_QUANTILE(value, 0.5) AS aq FROM groups GROUP BY category ORDER \ + BY category", + ); + let got_grouped: Vec> = wrk.read_stdout_on_success(&mut grouped_cmd); + assert_eq!( + got_grouped, + vec![svec!["category", "aq"], svec!["a", "2"], svec!["b", "5"]] + ); + + // Outside 2-4 arguments it is a syntax error, not a silent fallback. + let mut arity_cmd = wrk.command("sqlp"); + arity_cmd + .arg("groups.csv") + .arg("SELECT APPROX_QUANTILE(value) AS aq FROM groups"); + let stderr = wrk.stderr_on_error(&mut arity_cmd); + assert!( + stderr.contains("APPROX_QUANTILE expects 2-4 arguments (found 1)"), + "unexpected stderr: {stderr}" + ); +} + +/// Fixture for the window-function tests: two groups, correlated a/b pairs. +fn window_fixture(wrk: &Workdir) { + wrk.create( + "win.csv", + vec![ + svec!["i", "g", "a", "b"], + svec!["0", "a", "1", "1"], + svec!["1", "a", "2", "3"], + svec!["2", "a", "3", "2"], + svec!["3", "b", "4", "10"], + svec!["4", "b", "5", "20"], + ], + ); +} + +#[test] +fn sqlp_over_multi_arg_aggregates() { + // pola-rs/polars#29160: OVER now applies to aggregates that take more than + // one argument. Before the fix the window was dropped and these collapsed to + // a whole-frame aggregate. The existing OVER tests only cover the + // single-argument ranking functions (ROW_NUMBER/RANK/DENSE_RANK), so the + // per-partition values below are the regression signal. + let wrk = Workdir::new("sqlp_over_multi_arg_aggregates"); + window_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("win.csv") + .arg( + "SELECT i, g, CORR(a,b) OVER (PARTITION BY g) AS corr, COVAR_POP(a,b) OVER (PARTITION \ + BY g) AS cvp, COVAR_SAMP(a,b) OVER (PARTITION BY g) AS cvs, QUANTILE_CONT(a,0.5) \ + OVER (PARTITION BY g) AS qc, QUANTILE_DISC(a,0.5) OVER (PARTITION BY g) AS qd, \ + STRING_AGG(g,'-') OVER (PARTITION BY g) AS sa FROM win ORDER BY i", + ) + .args(["--float-precision", "6"]); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["i", "g", "corr", "cvp", "cvs", "qc", "qd", "sa"], + svec![ + "0", "a", "0.500000", "0.333333", "0.500000", "2.000000", "2.000000", "a-a-a" + ], + svec![ + "1", "a", "0.500000", "0.333333", "0.500000", "2.000000", "2.000000", "a-a-a" + ], + svec![ + "2", "a", "0.500000", "0.333333", "0.500000", "2.000000", "2.000000", "a-a-a" + ], + svec![ + "3", "b", "1.000000", "2.500000", "5.000000", "4.500000", "4.000000", "b-b" + ], + svec![ + "4", "b", "1.000000", "2.500000", "5.000000", "4.500000", "4.000000", "b-b" + ], + ]; + assert_eq!(got, expected); +} + +/// Fixture for the window NULLS ordering tests: nulls in both groups. +fn window_nulls_fixture(wrk: &Workdir) { + wrk.create( + "wnulls.csv", + vec![ + svec!["grp", "a"], + svec!["x", "20.0"], + svec!["x", ""], + svec!["x", "10.0"], + svec!["y", ""], + svec!["y", "40.0"], + svec!["y", "30.0"], + ], + ); +} + +#[test] +fn sqlp_window_order_by_nulls_last() { + // pola-rs/polars#29159: NULLS FIRST/LAST inside a *window's* ORDER BY is now + // respected. The existing NULLS FIRST/LAST tests in this file all sit on the + // top-level ORDER BY, which took a different code path. + let wrk = Workdir::new("sqlp_window_order_by_nulls_last"); + window_nulls_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("wnulls.csv").arg( + "SELECT grp, a, ROW_NUMBER() OVER (ORDER BY a NULLS LAST) AS rn, COUNT(*) OVER (ORDER BY \ + a NULLS LAST) AS cnt, SUM(a) OVER (ORDER BY a NULLS LAST) AS total FROM wnulls ORDER BY \ + rn", + ); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["grp", "a", "rn", "cnt", "total"], + svec!["x", "10.0", "1", "1", "10.0"], + svec!["x", "20.0", "2", "2", "30.0"], + svec!["y", "30.0", "3", "3", "60.0"], + svec!["y", "40.0", "4", "4", "100.0"], + svec!["x", "", "5", "5", "100.0"], + svec!["y", "", "6", "6", "100.0"], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_window_order_by_nulls_first() { + // pola-rs/polars#29159, the mirror of sqlp_window_order_by_nulls_last: the + // two nulls must lead, and DESC must not silently flip the null placement. + let wrk = Workdir::new("sqlp_window_order_by_nulls_first"); + window_nulls_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("wnulls.csv").arg( + "SELECT grp, a, ROW_NUMBER() OVER (ORDER BY a NULLS FIRST) AS rn FROM wnulls ORDER BY rn", + ); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["grp", "a", "rn"], + svec!["x", "", "1"], + svec!["y", "", "2"], + svec!["x", "10.0", "3"], + svec!["x", "20.0", "4"], + svec!["y", "30.0", "5"], + svec!["y", "40.0", "6"], + ] + ); + + let mut desc_cmd = wrk.command("sqlp"); + desc_cmd.arg("wnulls.csv").arg( + "SELECT grp, a, ROW_NUMBER() OVER (ORDER BY a DESC NULLS FIRST) AS rn FROM wnulls ORDER \ + BY rn", + ); + let got_desc: Vec> = wrk.read_stdout_on_success(&mut desc_cmd); + assert_eq!( + got_desc, + vec![ + svec!["grp", "a", "rn"], + svec!["x", "", "1"], + svec!["y", "", "2"], + svec!["y", "40.0", "3"], + svec!["y", "30.0", "4"], + svec!["x", "20.0", "5"], + svec!["x", "10.0", "6"], + ] + ); +} + +#[test] +fn sqlp_window_partition_by_order_by_nulls() { + // pola-rs/polars#29159 combined with PARTITION BY: the null sorts last + // *within* each partition, so both groups restart at rn 1. + let wrk = Workdir::new("sqlp_window_partition_by_order_by_nulls"); + window_nulls_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("wnulls.csv").arg( + "SELECT grp, a, ROW_NUMBER() OVER (PARTITION BY grp ORDER BY a NULLS LAST) AS rn FROM \ + wnulls ORDER BY grp, rn", + ); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["grp", "a", "rn"], + svec!["x", "10.0", "1"], + svec!["x", "20.0", "2"], + svec!["x", "", "3"], + svec!["y", "30.0", "1"], + svec!["y", "40.0", "2"], + svec!["y", "", "3"], + ] + ); +} + +#[test] +fn sqlp_over_aggregate_with_having() { + // pola-rs/polars#29006: an OVER window wrapping an aggregate coexists with + // HAVING -- the window sees the post-aggregation groups. + let wrk = Workdir::new("sqlp_over_aggregate_with_having"); + window_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("win.csv").arg( + "SELECT g, SUM(a) AS s, MAX(SUM(a)) OVER () AS mx FROM win GROUP BY g HAVING SUM(a) > 3 \ + ORDER BY g", + ); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["g", "s", "mx"], + svec!["a", "6", "9"], + svec!["b", "9", "9"], + ] + ); +} + +/// Fixture for the parenthesized-JOIN tests: overlapping a/b so that a +/// two-column constraint selects exactly one row. +fn paren_join_fixture(wrk: &Workdir) { + wrk.create( + "df1.csv", + vec![ + svec!["a", "b"], + svec!["1", "2"], + svec!["2", "3"], + svec!["3", "4"], + ], + ); + wrk.create( + "df2.csv", + vec![ + svec!["a", "b"], + svec!["2", "3"], + svec!["3", "9"], + svec!["9", "9"], + ], + ); +} + +#[test] +fn sqlp_join_parenthesized_constraint() { + // pola-rs/polars#28967: `ON ()` wrapped in parens is accepted. + // ORDER BY a1 is on a unique left key -- sqlp has no --maintain-order, so a + // join test without a total order is not repeatable. + let wrk = Workdir::new("sqlp_join_parenthesized_constraint"); + paren_join_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.args(["df1.csv", "df2.csv"]).arg( + "SELECT df1.a AS a1, df1.b AS b1, df2.a AS a2, df2.b AS b2 FROM df1 JOIN df2 ON (df1.a = \ + df2.a AND df1.b = df2.b) ORDER BY a1", + ); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![svec!["a1", "b1", "a2", "b2"], svec!["2", "3", "2", "3"],] + ); + + // LEFT JOIN keeps the non-matching left rows with nulls on the right. + let mut left_cmd = wrk.command("sqlp"); + left_cmd.args(["df1.csv", "df2.csv"]).arg( + "SELECT df1.a AS a1, df1.b AS b1, df2.a AS a2, df2.b AS b2 FROM df1 LEFT JOIN df2 ON \ + (df1.a = df2.a AND df1.b = df2.b) ORDER BY a1", + ); + let got_left: Vec> = wrk.read_stdout_on_success(&mut left_cmd); + assert_eq!( + got_left, + vec![ + svec!["a1", "b1", "a2", "b2"], + svec!["1", "2", "", ""], + svec!["2", "3", "2", "3"], + svec!["3", "4", "", ""], + ] + ); + + // A parenthesized non-predicate is rejected rather than treated as a cross + // join filter. + let mut bad_cmd = wrk.command("sqlp"); + bad_cmd + .args(["df1.csv", "df2.csv"]) + .arg("SELECT * FROM df1 JOIN df2 ON (df1.a)"); + let stderr = wrk.stderr_on_error(&mut bad_cmd); + assert!( + stderr.contains("predicates must resolve to boolean"), + "unexpected stderr: {stderr}" + ); +} + +#[test] +fn sqlp_join_parenthesized_relation_alias() { + // pola-rs/polars#29158: bare parens around a join do NOT introduce a new + // scope, so aliases declared inside them stay visible to the outer SELECT. + // Asserted by equivalence with the unparenthesized spelling. + let wrk = Workdir::new("sqlp_join_parenthesized_relation_alias"); + paren_join_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.args(["df1.csv", "df2.csv"]).arg( + "SELECT lhs.a, rhs.b FROM (df1 AS lhs INNER JOIN df2 AS rhs ON lhs.b = rhs.b) ORDER BY \ + lhs.a", + ); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![svec!["a", "b"], svec!["2", "3"]]; + assert_eq!(got, expected); + + let mut plain_cmd = wrk.command("sqlp"); + plain_cmd.args(["df1.csv", "df2.csv"]).arg( + "SELECT lhs.a, rhs.b FROM df1 AS lhs INNER JOIN df2 AS rhs ON lhs.b = rhs.b ORDER BY lhs.a", + ); + let got_plain: Vec> = wrk.read_stdout_on_success(&mut plain_cmd); + assert_eq!(got_plain, expected); +} + +#[test] +fn sqlp_join_literal_comparison() { + // pola-rs/polars#28701: a constant comparison in an ON clause belongs to the + // input it names, not to a post-join filter. The LEFT JOIN case is what + // distinguishes the two: `bob`/`charlie` fail `role = 'admin'`, so they must + // still appear with a null dept. A post-join filter would drop them. + let wrk = Workdir::new("sqlp_join_literal_comparison"); + wrk.create( + "people.csv", + vec![ + svec!["name", "role"], + svec!["alice", "admin"], + svec!["bob", "user"], + svec!["adam", "admin"], + svec!["charlie", "user"], + ], + ); + wrk.create( + "depts.csv", + vec![ + svec!["name", "dept"], + svec!["alice", "IT"], + svec!["bob", "HR"], + svec!["charlie", "IT"], + svec!["adam", "SEC"], + ], + ); + + let mut cmd = wrk.command("sqlp"); + cmd.args(["people.csv", "depts.csv"]).arg( + "SELECT people.name, people.role, depts.dept FROM people INNER JOIN depts ON people.name \ + = depts.name AND people.role = 'admin' ORDER BY people.name", + ); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["name", "role", "dept"], + svec!["adam", "admin", "SEC"], + svec!["alice", "admin", "IT"], + ] + ); + + let mut left_cmd = wrk.command("sqlp"); + left_cmd.args(["people.csv", "depts.csv"]).arg( + "SELECT people.name, people.role, depts.dept FROM people LEFT JOIN depts ON people.name = \ + depts.name AND people.role = 'admin' ORDER BY people.name", + ); + let got_left: Vec> = wrk.read_stdout_on_success(&mut left_cmd); + assert_eq!( + got_left, + vec![ + svec!["name", "role", "dept"], + svec!["adam", "admin", "SEC"], + svec!["alice", "admin", "IT"], + svec!["bob", "user", ""], + svec!["charlie", "user", ""], + ] + ); +} + +#[test] +fn sqlp_join_non_equi_decimal() { + // pola-rs/polars#29156 also taught the non-equi join path to compare + // Decimals. Reachable because qsv enables both `dtype-decimal` and `iejoin`. + // + // The values carry fractional parts ON PURPOSE: 1.05 < 1.10 is true in + // Decimal but false under any integer truncation, so the (1, 10) pair below + // is what proves the comparison really happens at decimal precision rather + // than the test passing on whole numbers that compare the same either way. + let wrk = Workdir::new("sqlp_join_non_equi_decimal"); + wrk.create( + "dec1.csv", + vec![ + svec!["id", "v"], + svec!["1", "1.05"], + svec!["2", "2.50"], + svec!["3", "10.10"], + ], + ); + wrk.create( + "dec2.csv", + vec![ + svec!["id", "w"], + svec!["10", "1.10"], + svec!["11", "2.50"], + svec!["12", "9.99"], + ], + ); + + let mut cmd = wrk.command("sqlp"); + cmd.args(["dec1.csv", "dec2.csv"]).arg( + "SELECT dec1.id AS l, dec1.v, dec2.id AS r, dec2.w FROM dec1 JOIN dec2 ON CAST(dec1.v AS \ + DECIMAL(10,2)) < CAST(dec2.w AS DECIMAL(10,2)) ORDER BY l, r", + ); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["l", "v", "r", "w"], + svec!["1", "1.05", "10", "1.1"], + svec!["1", "1.05", "11", "2.5"], + svec!["1", "1.05", "12", "9.99"], + svec!["2", "2.5", "12", "9.99"], + ] + ); + + // >= pins exact boundary equality: 2.50 vs 2.50 matches, and 10.10 exceeds + // every right-hand value. + let mut ge_cmd = wrk.command("sqlp"); + ge_cmd.args(["dec1.csv", "dec2.csv"]).arg( + "SELECT dec1.id AS l, dec2.id AS r FROM dec1 JOIN dec2 ON CAST(dec1.v AS DECIMAL(10,2)) \ + >= CAST(dec2.w AS DECIMAL(10,2)) ORDER BY l, r", + ); + let got_ge: Vec> = wrk.read_stdout_on_success(&mut ge_cmd); + assert_eq!( + got_ge, + vec![ + svec!["l", "r"], + svec!["2", "10"], + svec!["2", "11"], + svec!["3", "10"], + svec!["3", "11"], + svec!["3", "12"], + ] + ); +} + +/// Fixture for the subquery / scope tests. `t2` deliberately names its columns +/// `b, a` so an unqualified reference would be ambiguous across the two. +fn subquery_fixture(wrk: &Workdir) { + wrk.create( + "t1.csv", + vec![ + svec!["a", "b"], + svec!["1", "10"], + svec!["2", "20"], + svec!["3", "30"], + ], + ); + wrk.create( + "t2.csv", + vec![svec!["b", "a"], svec!["1", "100"], svec!["2", "200"]], + ); +} + +#[test] +fn sqlp_exists_subquery() { + // pola-rs/polars#29006: EXISTS / NOT EXISTS correlated on the outer relation. + // Nothing in this file covered EXISTS before -- only derived-table subqueries. + let wrk = Workdir::new("sqlp_exists_subquery"); + subquery_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.args(["t1.csv", "t2.csv"]) + .arg("SELECT a, b FROM t1 WHERE EXISTS (SELECT 1 FROM t2 WHERE t2.b = t1.a) ORDER BY a"); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![svec!["a", "b"], svec!["1", "10"], svec!["2", "20"]] + ); + + let mut not_cmd = wrk.command("sqlp"); + not_cmd.args(["t1.csv", "t2.csv"]).arg( + "SELECT a, b FROM t1 WHERE NOT EXISTS (SELECT 1 FROM t2 WHERE t2.b = t1.a) ORDER BY a", + ); + let got_not: Vec> = wrk.read_stdout_on_success(&mut not_cmd); + assert_eq!(got_not, vec![svec!["a", "b"], svec!["3", "30"]]); +} + +#[test] +fn sqlp_correlated_scalar_subquery() { + // pola-rs/polars#28939: a scalar subquery's aggregate binds to the + // subquery's OWN relation. `SUM(x.a)` therefore sums t2's `a`, and a row + // with no match yields null rather than 0. + let wrk = Workdir::new("sqlp_correlated_scalar_subquery"); + subquery_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.args(["t1.csv", "t2.csv"]) + .arg("SELECT a, (SELECT SUM(x.a) FROM t2 x WHERE x.b = t1.a) AS s FROM t1 ORDER BY a"); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["a", "s"], + svec!["1", "100"], + svec!["2", "200"], + svec!["3", ""], + ] + ); + + // Aggregating a column of the OUTER relation inside the subquery is now an + // error instead of silently resolving. + let mut bad_cmd = wrk.command("sqlp"); + bad_cmd + .args(["t1.csv", "t2.csv"]) + .arg("SELECT a, (SELECT SUM(t1.b) FROM t2 x WHERE x.b = t1.a) AS s FROM t1"); + let stderr = wrk.stderr_on_error(&mut bad_cmd); + assert!( + stderr.contains("no table or struct column named 't1' found"), + "unexpected stderr: {stderr}" + ); +} + +#[test] +fn sqlp_scope_strictness() { + // pola-rs/polars#28937: SQL scope rules now follow Postgres. Once a relation + // is aliased, its original name is out of scope -- but an unqualified + // ORDER BY key still resolves, and an outer alias is still visible to a + // correlated subquery. + let wrk = Workdir::new("sqlp_scope_strictness"); + subquery_fixture(&wrk); + + // The alias hides the table name. + let mut bad_cmd = wrk.command("sqlp"); + bad_cmd.arg("t1.csv").arg("SELECT t1.a FROM t1 AS f"); + let stderr = wrk.stderr_on_error(&mut bad_cmd); + assert!( + stderr.contains("no table or struct column named 't1' found"), + "unexpected stderr: {stderr}" + ); + + let ok_queries = [ + // qualified by the alias, ordered by the bare column name + "SELECT f.a FROM t1 AS f ORDER BY a", + // unaliased, so the table name is still in scope + "SELECT t1.a FROM t1 ORDER BY a", + // the outer alias `o` is visible inside the correlated IN subquery + "SELECT a FROM t1 o WHERE b IN (SELECT x.b FROM t1 AS x WHERE x.a <= o.a) ORDER BY a", + ]; + for query in ok_queries { + let mut cmd = wrk.command("sqlp"); + cmd.arg("t1.csv").arg(query); + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![svec!["a"], svec!["1"], svec!["2"], svec!["3"]], + "query: {query}" + ); + } +} + +#[test] +fn sqlp_group_by_unaliased_constants() { + // pola-rs/polars#29367: unaliased constants in a SELECT with GROUP BY are + // projected once per group and named `literal`, `literal:1`, `literal:2` + // rather than colliding on a single `literal` column. + let wrk = Workdir::new("sqlp_group_by_unaliased_constants"); + subquery_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("t1.csv") + .arg("SELECT 2, a, 'x', COUNT(*) AS n, 3 FROM t1 GROUP BY a ORDER BY a"); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + assert_eq!( + got, + vec![ + svec!["literal", "a", "literal:1", "n", "literal:2"], + svec!["2", "1", "x", "1", "3"], + svec!["2", "2", "x", "1", "3"], + svec!["2", "3", "x", "1", "3"], + ] + ); +} + +#[test] +fn sqlp_order_by_aggregate() { + // pola-rs/polars#29010: ORDER BY accepts a restated aggregate, an aggregate + // that is not in the SELECT list at all, COUNT(*), and an aggregate + // expression. `b` outranks `a` on SUM(value) (15 vs 6) while the two tie on + // COUNT(*), so the COUNT case needs `category` as a tiebreak to be a total + // order. + let wrk = Workdir::new("sqlp_order_by_aggregate"); + grouping_fixture(&wrk); + + let by_total = vec![ + svec!["category", "total"], + svec!["b", "15"], + svec!["a", "6"], + ]; + + let mut restated_cmd = wrk.command("sqlp"); + restated_cmd.arg("groups.csv").arg( + "SELECT category, SUM(value) AS total FROM groups GROUP BY category ORDER BY SUM(value) \ + DESC", + ); + let got_restated: Vec> = wrk.read_stdout_on_success(&mut restated_cmd); + assert_eq!(got_restated, by_total); + + // ORDER BY an aggregate expression, not just the bare aggregate. + let mut expr_cmd = wrk.command("sqlp"); + expr_cmd.arg("groups.csv").arg( + "SELECT category, SUM(value) AS total FROM groups GROUP BY category ORDER BY SUM(value) * \ + -1", + ); + let got_expr: Vec> = wrk.read_stdout_on_success(&mut expr_cmd); + assert_eq!(got_expr, by_total); + + // The aggregate need not be projected. + let mut unselected_cmd = wrk.command("sqlp"); + unselected_cmd + .arg("groups.csv") + .arg("SELECT category FROM groups GROUP BY category ORDER BY SUM(value) DESC"); + let got_unselected: Vec> = wrk.read_stdout_on_success(&mut unselected_cmd); + assert_eq!( + got_unselected, + vec![svec!["category"], svec!["b"], svec!["a"]] + ); + + // COUNT(*) ties at 3 for both groups, so `category` decides. + let mut count_cmd = wrk.command("sqlp"); + count_cmd + .arg("groups.csv") + .arg("SELECT category FROM groups GROUP BY category ORDER BY COUNT(*) DESC, category"); + let got_count: Vec> = wrk.read_stdout_on_success(&mut count_cmd); + assert_eq!(got_count, vec![svec!["category"], svec!["a"], svec!["b"]]); +} + +#[test] +fn sqlp_cte_shadows_table_and_case_insensitive_relation() { + // pola-rs/polars#29006: a CTE shadows a same-named registered table, and a + // relation name resolves case-insensitively. + let wrk = Workdir::new("sqlp_cte_shadows_table_and_case_insensitive_relation"); + grouping_fixture(&wrk); + + let mut cte_cmd = wrk.command("sqlp"); + cte_cmd.arg("groups.csv").arg( + "WITH groups AS (SELECT 'z' AS category, 99 AS value) SELECT category, value FROM groups", + ); + let got_cte: Vec> = wrk.read_stdout_on_success(&mut cte_cmd); + assert_eq!(got_cte, vec![svec!["category", "value"], svec!["z", "99"]]); + + let mut case_cmd = wrk.command("sqlp"); + case_cmd + .arg("groups.csv") + .arg("SELECT category FROM GROUPS GROUP BY category ORDER BY category"); + let got_case: Vec> = wrk.read_stdout_on_success(&mut case_cmd); + assert_eq!(got_case, vec![svec!["category"], svec!["a"], svec!["b"]]); +} + +#[test] +fn sqlp_date_part_functions() { + // pola-rs/polars#29269 added the bare date-part shorthands to the SQL + // frontend: YEAR, QUARTER, MONTH, WEEK, DAY/DAYOFMONTH, DAYOFWEEK, + // DAYOFYEAR, HOUR, MINUTE, SECOND. None of them existed at py-1.44.0. + // + // The 2024-12-31 row is the discriminating one: WEEK is ISO week, so it is + // **1** (of the following year), not 53 as a naive week-of-year would give. + // DAYOFWEEK is ISO too -- Monday is 1, so the Monday 2021-03-15 is 1 and the + // Tuesday 2024-12-31 is 2. DAYOFYEAR 366 confirms the leap year. + let wrk = Workdir::new("sqlp_date_part_functions"); + wrk.create( + "dtparts.csv", + vec![ + svec!["ts", "d"], + svec!["2021-03-15 10:30:20", "2021-03-15"], + svec!["2024-12-31 23:59:59", "2024-12-31"], + ], + ); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("dtparts.csv") + .arg( + "SELECT ts, YEAR(ts) AS y, QUARTER(ts) AS q, MONTH(ts) AS mo, WEEK(ts) AS wk, DAY(ts) \ + AS d, DAYOFMONTH(ts) AS dom, DAYOFWEEK(ts) AS dow, DAYOFYEAR(ts) AS doy, HOUR(ts) AS \ + h, MINUTE(ts) AS mi, SECOND(ts) AS s FROM dtparts", + ) + .arg("--try-parsedates"); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec![ + "ts", "y", "q", "mo", "wk", "d", "dom", "dow", "doy", "h", "mi", "s" + ], + svec![ + "2021-03-15T10:30:20.000000", + "2021", + "1", + "3", + "11", + "15", + "15", + "1", + "74", + "10", + "30", + "20" + ], + svec![ + "2024-12-31T23:59:59.000000", + "2024", + "4", + "12", + "1", + "31", + "31", + "2", + "366", + "23", + "59", + "59" + ], + ]; + assert_eq!(got, expected); + + // They apply to a Date column too. `d` is date-shaped, so --try-parsedates + // makes it a real Date; note CAST(ts AS DATE) would NOT work here, because + // format inference cannot read a datetime string as a date. + let mut date_cmd = wrk.command("sqlp"); + date_cmd + .arg("dtparts.csv") + .arg("SELECT YEAR(d) AS y, MONTH(d) AS mo, DAY(d) AS dd, WEEK(d) AS wk FROM dtparts") + .arg("--try-parsedates"); + let got_date: Vec> = wrk.read_stdout_on_success(&mut date_cmd); + assert_eq!( + got_date, + vec![ + svec!["y", "mo", "dd", "wk"], + svec!["2021", "3", "15", "11"], + svec!["2024", "12", "31", "1"], + ] + ); +} + +#[test] +fn sqlp_grouping_id() { + // pola-rs/polars#29278 added GROUPING_ID alongside GROUPING(). On a ROLLUP + // the two agree: both pack one bit per argument, MSB-first, so the grand + // total is 3 (both keys rolled up) and a category subtotal is 1 (only class + // rolled up). + let wrk = Workdir::new("sqlp_grouping_id"); + grouping_fixture(&wrk); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("groups.csv") + .arg( + "SELECT category, class, SUM(value) AS total, GROUPING_ID(category, class) AS gid, \ + GROUPING(category, class) AS g FROM groups GROUP BY ROLLUP(category, class) ORDER BY \ + category NULLS LAST, class NULLS LAST", + ) + .args(["--wnull-value", "NULL"]); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["category", "class", "total", "gid", "g"], + svec!["a", "x", "3", "0", "0"], + svec!["a", "y", "3", "0", "0"], + svec!["a", "NULL", "6", "1", "1"], + svec!["b", "x", "4", "0", "0"], + svec!["b", "y", "11", "0", "0"], + svec!["b", "NULL", "15", "1", "1"], + svec!["NULL", "NULL", "21", "3", "3"], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_cast_temporal_error_payloads() { + // pola-rs/polars#28986 routes SQL string->temporal casts through + // strptime-infer, which reports TWO different failures depending on whether + // a format could be inferred at all: + // * some rows parse -> InvalidOperation, "conversion ... failed", naming the offending + // value (see sqlp_cast_strict_vs_try_temporal) + // * NO row parses -> ComputeError, "could not find an appropriate format to parse + // " + // Before the rewrite both cases produced the first message, so the + // format-inference failure is the new signal. + let wrk = Workdir::new("sqlp_cast_temporal_error_payloads"); + wrk.create( + "allbad.csv", + vec![svec!["id", "s"], svec!["1", "aaa"], svec!["2", "bbb"]], + ); + + let mut date_cmd = wrk.command("sqlp"); + date_cmd + .arg("allbad.csv") + .arg("SELECT id, CAST(s AS DATE) AS d FROM allbad"); + let date_stderr = wrk.stderr_on_error(&mut date_cmd); + assert!( + date_stderr.contains("could not find an appropriate format to parse dates"), + "unexpected stderr: {date_stderr}" + ); + + let mut time_cmd = wrk.command("sqlp"); + time_cmd + .arg("allbad.csv") + .arg("SELECT id, CAST(s AS TIME) AS t FROM allbad"); + let time_stderr = wrk.stderr_on_error(&mut time_cmd); + assert!( + time_stderr.contains("could not find an appropriate format to parse times"), + "unexpected stderr: {time_stderr}" + ); + + // TRY_CAST still degrades to all-null rather than failing. + let mut try_cmd = wrk.command("sqlp"); + try_cmd + .arg("allbad.csv") + .arg("SELECT id, TRY_CAST(s AS DATE) AS d FROM allbad ORDER BY id"); + let got: Vec> = wrk.read_stdout_on_success(&mut try_cmd); + assert_eq!(got, vec![svec!["id", "d"], svec!["1", ""], svec!["2", ""]]); +} From a4e842e588b8262d33270f0fd3f55207a7146dae Mon Sep 17 00:00:00 2001 From: Joel Natividad <1980690+jqnatividad@users.noreply.github.com> Date: Sat, 19 Sep 2026 13:46:18 -0400 Subject: [PATCH 2/4] test: make the window NULLS tests order-safe, and cover a real NULL grouping key roborev 4815, both findings LOW and both valid. Finding 1 - the two null rows in the #29159 window tests are peers, so asserting `x` before `y` relied on an order SQL does not define. Correct, and the same species of latent fragility the repo already fixed once in 6d33fb41f ("it held over 10-12 local runs, which is not a guarantee"). I could not make the order swap - stable across POLARS_MAX_THREADS 1/2/4/8/16 and with 40 tied null rows - so this was a latent risk rather than an active flake, but the streaming engine is now the default for collect and the tie is free to move. The suggested fix - add `grp` as a secondary window ORDER BY key - is NOT usable, which is the interesting part: * `OVER (ORDER BY a NULLS FIRST, grp)` -> hard error, "OVER does not (yet) support mixed NULLS FIRST/LAST ordering for ORDER BY" * `OVER (ORDER BY a DESC NULLS FIRST, grp)` -> hard error, "OVER does not (yet) support mixed asc/desc directions for ORDER BY" * `OVER (ORDER BY a NULLS LAST, grp)` -> accepted, and SILENTLY puts the nulls FIRST. Also with `grp NULLS LAST` spelled out. So a window ORDER BY honors NULLS FIRST/LAST only with a single key; a second key drops the clause. The identical two-key ORDER BY at the TOP level honors it, so this is specific to the window path - an upstream gap in #29159 itself. Applying the suggestion would therefore have inverted the very property under test. Instead the two global-window tests stop projecting `grp`: `a`, `rn`, `cnt` and `total` are identical whichever peer lands at rn 5 vs 6, so the assertion is a total order over what it actually asserts. The PARTITION BY test keeps `grp` and its exact order - one null per partition means no peers - with a comment saying why it may. Adds sqlp_window_order_by_multiple_keys_limitation to pin all of the above, so the workaround becomes available again the moment upstream fixes it. Finding 2 - grouping_fixture has no NULL in the data, so nothing exercised the distinction GROUPING()/GROUPING_ID() exist for. Correct. Adds sqlp_grouping_distinguishes_a_data_null with real NULL categories, where two output rows both print NULL in `category` and only GROUPING() separates them: g=0 is the genuine NULL group (total 12), g=1 is the ROLLUP grand total (total 15). Added as its own test rather than by editing grouping_fixture, which nine existing tests depend on. Verified: full suite 3,960 passed / 0 failed; cargo t sqlp 5x green (136); 5 mutation tests on the new and reworked assertions, all 5 caught; double-run-check clean. Co-Authored-By: Claude Opus 5 (1M context) --- tests/test_sqlp.rs | 215 +++++++++++++++++++++++++++++++++++++++------ 1 file changed, 186 insertions(+), 29 deletions(-) diff --git a/tests/test_sqlp.rs b/tests/test_sqlp.rs index 8fef187aeb..1884115eed 100644 --- a/tests/test_sqlp.rs +++ b/tests/test_sqlp.rs @@ -5181,22 +5181,29 @@ fn sqlp_window_order_by_nulls_last() { let wrk = Workdir::new("sqlp_window_order_by_nulls_last"); window_nulls_fixture(&wrk); + // `grp` is deliberately NOT projected. The two null-`a` rows are peers in + // this window, so which one gets rn 5 vs rn 6 is unspecified -- but `a`, + // `rn`, `cnt` and `total` are identical either way, so the assertion is a + // total order on what it actually asserts. Adding a second window ORDER BY + // key as a tiebreak is NOT an option: polars silently ignores + // NULLS FIRST/LAST once a window has more than one key (pinned by + // sqlp_window_order_by_multiple_keys_limitation), which would invert the + // very property under test. let mut cmd = wrk.command("sqlp"); cmd.arg("wnulls.csv").arg( - "SELECT grp, a, ROW_NUMBER() OVER (ORDER BY a NULLS LAST) AS rn, COUNT(*) OVER (ORDER BY \ - a NULLS LAST) AS cnt, SUM(a) OVER (ORDER BY a NULLS LAST) AS total FROM wnulls ORDER BY \ - rn", + "SELECT a, ROW_NUMBER() OVER (ORDER BY a NULLS LAST) AS rn, COUNT(*) OVER (ORDER BY a \ + NULLS LAST) AS cnt, SUM(a) OVER (ORDER BY a NULLS LAST) AS total FROM wnulls ORDER BY rn", ); let got: Vec> = wrk.read_stdout_on_success(&mut cmd); let expected = vec![ - svec!["grp", "a", "rn", "cnt", "total"], - svec!["x", "10.0", "1", "1", "10.0"], - svec!["x", "20.0", "2", "2", "30.0"], - svec!["y", "30.0", "3", "3", "60.0"], - svec!["y", "40.0", "4", "4", "100.0"], - svec!["x", "", "5", "5", "100.0"], - svec!["y", "", "6", "6", "100.0"], + svec!["a", "rn", "cnt", "total"], + svec!["10.0", "1", "1", "10.0"], + svec!["20.0", "2", "2", "30.0"], + svec!["30.0", "3", "3", "60.0"], + svec!["40.0", "4", "4", "100.0"], + svec!["", "5", "5", "100.0"], + svec!["", "6", "6", "100.0"], ]; assert_eq!(got, expected); } @@ -5208,40 +5215,40 @@ fn sqlp_window_order_by_nulls_first() { let wrk = Workdir::new("sqlp_window_order_by_nulls_first"); window_nulls_fixture(&wrk); + // As in sqlp_window_order_by_nulls_last, `grp` is left out: the two null + // rows are peers, so only `a` and `rn` are determinate. let mut cmd = wrk.command("sqlp"); - cmd.arg("wnulls.csv").arg( - "SELECT grp, a, ROW_NUMBER() OVER (ORDER BY a NULLS FIRST) AS rn FROM wnulls ORDER BY rn", - ); + cmd.arg("wnulls.csv") + .arg("SELECT a, ROW_NUMBER() OVER (ORDER BY a NULLS FIRST) AS rn FROM wnulls ORDER BY rn"); let got: Vec> = wrk.read_stdout_on_success(&mut cmd); assert_eq!( got, vec![ - svec!["grp", "a", "rn"], - svec!["x", "", "1"], - svec!["y", "", "2"], - svec!["x", "10.0", "3"], - svec!["x", "20.0", "4"], - svec!["y", "30.0", "5"], - svec!["y", "40.0", "6"], + svec!["a", "rn"], + svec!["", "1"], + svec!["", "2"], + svec!["10.0", "3"], + svec!["20.0", "4"], + svec!["30.0", "5"], + svec!["40.0", "6"], ] ); let mut desc_cmd = wrk.command("sqlp"); desc_cmd.arg("wnulls.csv").arg( - "SELECT grp, a, ROW_NUMBER() OVER (ORDER BY a DESC NULLS FIRST) AS rn FROM wnulls ORDER \ - BY rn", + "SELECT a, ROW_NUMBER() OVER (ORDER BY a DESC NULLS FIRST) AS rn FROM wnulls ORDER BY rn", ); let got_desc: Vec> = wrk.read_stdout_on_success(&mut desc_cmd); assert_eq!( got_desc, vec![ - svec!["grp", "a", "rn"], - svec!["x", "", "1"], - svec!["y", "", "2"], - svec!["y", "40.0", "3"], - svec!["y", "30.0", "4"], - svec!["x", "20.0", "5"], - svec!["x", "10.0", "6"], + svec!["a", "rn"], + svec!["", "1"], + svec!["", "2"], + svec!["40.0", "3"], + svec!["30.0", "4"], + svec!["20.0", "5"], + svec!["10.0", "6"], ] ); } @@ -5250,6 +5257,10 @@ fn sqlp_window_order_by_nulls_first() { fn sqlp_window_partition_by_order_by_nulls() { // pola-rs/polars#29159 combined with PARTITION BY: the null sorts last // *within* each partition, so both groups restart at rn 1. + // + // Unlike the two tests above this one CAN project `grp` and assert an exact + // order: each partition holds exactly one null, so there are no peers and + // every row's rn is determined. let wrk = Workdir::new("sqlp_window_partition_by_order_by_nulls"); window_nulls_fixture(&wrk); @@ -5899,3 +5910,149 @@ fn sqlp_cast_temporal_error_payloads() { let got: Vec> = wrk.read_stdout_on_success(&mut try_cmd); assert_eq!(got, vec![svec!["id", "d"], svec!["1", ""], svec!["2", ""]]); } + +#[test] +fn sqlp_grouping_distinguishes_a_data_null() { + // pola-rs/polars#29278: this is the whole reason GROUPING()/GROUPING_ID() + // exist. A ROLLUP writes NULL into the keys it rolls up, which is + // indistinguishable *in the data* from a group whose key is genuinely NULL. + // The fixture therefore carries real NULL categories (empty CSV fields read + // as NULL), so two output rows both print NULL in `category` and only + // GROUPING() tells them apart: + // category=NULL, g=0 -> the real NULL group, total 12 (5 + 7) + // category=NULL, g=1 -> the ROLLUP grand total, total 15 (3 + 12) + // The shared `grouping_fixture` has no NULL in the data and so cannot reach + // this distinction at all. + let wrk = Workdir::new("sqlp_grouping_distinguishes_a_data_null"); + wrk.create( + "gnull.csv", + vec![ + svec!["category", "class", "value"], + svec!["a", "x", "1"], + svec!["a", "y", "2"], + svec!["", "x", "5"], + svec!["", "y", "7"], + ], + ); + + let mut cmd = wrk.command("sqlp"); + cmd.arg("gnull.csv") + .arg( + "SELECT category, SUM(value) AS total, GROUPING(category) AS g FROM gnull GROUP BY \ + ROLLUP(category) ORDER BY g, category NULLS LAST", + ) + .args(["--wnull-value", "NULL"]); + + let got: Vec> = wrk.read_stdout_on_success(&mut cmd); + let expected = vec![ + svec!["category", "total", "g"], + svec!["a", "3", "0"], + // a REAL null key: GROUPING is 0, because `category` participates in + // this grouping set -- the NULL is data, not a subtotal marker. + svec!["NULL", "12", "0"], + // the rolled-up grand total: same printed NULL, GROUPING is 1. + svec!["NULL", "15", "1"], + ]; + assert_eq!(got, expected); +} + +#[test] +fn sqlp_window_order_by_multiple_keys_limitation() { + // Found while hardening the #29159 tests, and the reason they project only + // determinate columns instead of adding a tiebreak key. + // + // A window ORDER BY honors NULLS FIRST/LAST only with a SINGLE key. Add a + // second key and the NULLS clause is SILENTLY IGNORED -- the nulls move to + // the front even when every key says NULLS LAST. A top-level ORDER BY with + // two keys honors it correctly, so this is specific to the window path. + // + // Mixing directions or NULLS placement across window keys is rejected + // outright. If a future polars fixes any of this, these assertions fail and + // the tiebreak workaround becomes available. + let wrk = Workdir::new("sqlp_window_order_by_multiple_keys_limitation"); + window_nulls_fixture(&wrk); + + // Single key: NULLS LAST is honored -- nulls get rn 5 and 6. + let mut single_cmd = wrk.command("sqlp"); + single_cmd + .arg("wnulls.csv") + .arg("SELECT a, ROW_NUMBER() OVER (ORDER BY a NULLS LAST) AS rn FROM wnulls ORDER BY rn"); + let got_single: Vec> = wrk.read_stdout_on_success(&mut single_cmd); + assert_eq!( + got_single, + vec![ + svec!["a", "rn"], + svec!["10.0", "1"], + svec!["20.0", "2"], + svec!["30.0", "3"], + svec!["40.0", "4"], + svec!["", "5"], + svec!["", "6"], + ] + ); + + // Two keys, BOTH spelled NULLS LAST: the clause is dropped and the nulls + // lead. This is the bug -- it is pinned, not endorsed. + let mut multi_cmd = wrk.command("sqlp"); + multi_cmd.arg("wnulls.csv").arg( + "SELECT grp, a, ROW_NUMBER() OVER (ORDER BY a NULLS LAST, grp NULLS LAST) AS rn FROM \ + wnulls ORDER BY rn", + ); + let got_multi: Vec> = wrk.read_stdout_on_success(&mut multi_cmd); + assert_eq!( + got_multi, + vec![ + svec!["grp", "a", "rn"], + svec!["x", "", "1"], + svec!["y", "", "2"], + svec!["x", "10.0", "3"], + svec!["x", "20.0", "4"], + svec!["y", "30.0", "5"], + svec!["y", "40.0", "6"], + ] + ); + + // The same two-key ORDER BY at the TOP level honors NULLS LAST, which is + // what makes the above a window-specific defect rather than a syntax quirk. + let mut top_cmd = wrk.command("sqlp"); + top_cmd + .arg("wnulls.csv") + .arg("SELECT grp, a FROM wnulls ORDER BY a NULLS LAST, grp"); + let got_top: Vec> = wrk.read_stdout_on_success(&mut top_cmd); + assert_eq!( + got_top, + vec![ + svec!["grp", "a"], + svec!["x", "10.0"], + svec!["x", "20.0"], + svec!["y", "30.0"], + svec!["y", "40.0"], + svec!["x", ""], + svec!["y", ""], + ] + ); + + // Mixed NULLS placement across window keys is a hard error. + let mut mixed_nulls_cmd = wrk.command("sqlp"); + mixed_nulls_cmd.arg("wnulls.csv").arg( + "SELECT a, ROW_NUMBER() OVER (ORDER BY a NULLS FIRST, grp) AS rn FROM wnulls ORDER BY rn", + ); + let mixed_nulls_stderr = wrk.stderr_on_error(&mut mixed_nulls_cmd); + assert!( + mixed_nulls_stderr + .contains("OVER does not (yet) support mixed NULLS FIRST/LAST ordering for ORDER BY"), + "unexpected stderr: {mixed_nulls_stderr}" + ); + + // So is a mixed asc/desc window ORDER BY. + let mut mixed_dir_cmd = wrk.command("sqlp"); + mixed_dir_cmd.arg("wnulls.csv").arg( + "SELECT a, ROW_NUMBER() OVER (ORDER BY a DESC NULLS FIRST, grp) AS rn FROM wnulls ORDER \ + BY rn", + ); + let mixed_dir_stderr = wrk.stderr_on_error(&mut mixed_dir_cmd); + assert!( + mixed_dir_stderr.contains("OVER does not (yet) support mixed asc/desc directions"), + "unexpected stderr: {mixed_dir_stderr}" + ); +} From 38b30c6b6fb4492466dacec94bb191dcc14607a8 Mon Sep 17 00:00:00 2001 From: Joel Natividad <1980690+jqnatividad@users.noreply.github.com> Date: Sat, 19 Sep 2026 13:51:56 -0400 Subject: [PATCH 3/4] docs: cite pola-rs/polars#29390 and scope the window NULLS claim to one key Filed the multi-key window ORDER BY defect found while addressing roborev 4815 as pola-rs/polars#29390, with the root cause: Expr::over_with_options (polars-plan/src/dsl/mod.rs:853-861) collapses several ORDER BY keys into a single as_struct(...), and a struct holding a null field is not itself null, so SortOptions::nulls_last has nothing to act on. `descending` survives the same path, which is why only the null placement is wrong. polars-sql's own parse_order_by_in_window is correct - it validates uniformity and derives one SortOptions - so the option is lost below the SQL layer. Measured: with two keys the null placement tracks only the direction (nulls first under ASC, last under DESC) no matter what the NULLS clause asks for, so 2 of the 4 direction x placement combinations are wrong and the other 2 agree only by coincidence. Single key is correct in all 4. - sqlp_window_order_by_multiple_keys_limitation now cites the issue and the root cause, so the next reader can check whether it is fixed rather than rediscover it. - The CHANGELOG bullet claimed NULLS FIRST/LAST in a window ORDER BY works, full stop. Scoped to a single sort key and points at the upstream issue - a reader would otherwise reasonably assume it holds with a tiebreak key, which is the exact trap. Co-Authored-By: Claude Opus 5 (1M context) --- CHANGELOG.md | 2 +- tests/test_sqlp.rs | 6 ++++++ 2 files changed, 7 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 46f885f772..fb8cfc9619 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **the bare date-part functions `YEAR`, `QUARTER`, `MONTH`, `WEEK`, `DAY`/`DAYOFMONTH`, `DAYOFWEEK`, `DAYOFYEAR`, `HOUR`, `MINUTE`, `SECOND`** ([#29269](https://github.com/pola-rs/polars/pull/29269)). Both are ISO-based, which is the part worth knowing: `WEEK('2024-12-31')` is **1** (ISO week 1 of 2025), not 53, and `DAYOFWEEK` counts Monday as 1. - **`DATE '...'` / `TIMESTAMP '...'` typed literals** ([#29007](https://github.com/pola-rs/polars/pull/29007)), and **date +/- integer arithmetic** shifting by whole days with the integer on either side, from a literal or a column, plus **Decimal comparisons in non-equi joins** ([#29156](https://github.com/pola-rs/polars/pull/29156)). - **parenthesized `JOIN ... ON (...)` constraints** ([#28967](https://github.com/pola-rs/polars/pull/28967)); relation aliases declared *inside* those parens now resolve in the outer SELECT, since bare parens around a join do not open a new scope ([#29158](https://github.com/pola-rs/polars/pull/29158)). - - **`OVER` on multi-argument aggregates** - `CORR`, `COVAR_POP`, `COVAR_SAMP`, `QUANTILE_CONT`, `QUANTILE_DISC`, `STRING_AGG` ([#29160](https://github.com/pola-rs/polars/pull/29160)) - and **`NULLS FIRST`/`NULLS LAST` inside a window's own `ORDER BY`** ([#29159](https://github.com/pola-rs/polars/pull/29159)). + - **`OVER` on multi-argument aggregates** - `CORR`, `COVAR_POP`, `COVAR_SAMP`, `QUANTILE_CONT`, `QUANTILE_DISC`, `STRING_AGG` ([#29160](https://github.com/pola-rs/polars/pull/29160)) - and **`NULLS FIRST`/`NULLS LAST` inside a window's own `ORDER BY`** ([#29159](https://github.com/pola-rs/polars/pull/29159)) - though only when the window has a **single** sort key: with two or more keys the `NULLS` clause is silently ignored and the nulls follow the sort direction instead, which we reported upstream as [#29390](https://github.com/pola-rs/polars/issues/29390). - **`EXISTS`/`NOT EXISTS` and correlated scalar subqueries**, CTEs that shadow a registered table, case-insensitive relation names, and `ORDER BY` over an aggregate that is not in the SELECT list ([#29006](https://github.com/pola-rs/polars/pull/29006), [#29010](https://github.com/pola-rs/polars/pull/29010)). - `sqlp`: **`APPROX_QUANTILE(col, q [, allowed_rank_error [, method]])`** - an approximate quantile aggregate, new to the Polars SQL frontend in [pola-rs/polars#29288](https://github.com/pola-rs/polars/pull/29288). Upstream cfg-gates it behind a `approx_quantile` feature that ships only inside their `docs-selection`/`full` umbrellas, so qsv now enables it explicitly - without that, `sqlp` rejected the function as `unsupported function 'approx_quantile'`. `allowed_rank_error` defaults to `0.001` and `method` selects the sketch (e.g. `'kll'`); anything outside 2-4 arguments is a syntax error rather than a silent fallback. Being sketch-based it trades exactness for speed on large inputs - use `QUANTILE_CONT`/`QUANTILE_DISC` when you need the exact value. - **gallery: a real per-capita county rate map where the rate does NOT invert the ranking.** The gallery's two existing `--denominator census` figures both show a raw-count choropleth whose ranking *inverts* once divided by population - `district_requests` on six hand-drawn rectangles, `tristate_county_incidents` on 210 real counties with counts constructed to be anti-correlated with population. Both teach the lesson well and both are synthetic in exactly the place that matters, so a reader could fairly conclude that per-capita always flips a map. `wpa_211_requests` is the counter-example, on data that is real all the way down: **279,464 2-1-1 helpline requests** across the **25 Western Pennsylvania counties** served by the United Way of Southwestern Pennsylvania - the upstream resource *entire*, 2023-09-22 to 2025-08-27, no window and no sample (via the [WPRDC](https://data.wprdc.org/dataset/211-requests)). The count panel is Allegheny County and almost nothing else - **172,035 requests, 61.6% of the file** - and the rate panel beside it leaves Allegheny on top at **138.9 per 1,000 residents**. What moves is the middle: **Venango climbs 11th to 3rd**, **Cambria 5th to 2nd**, **Butler falls 8th to 13th**, and the two top-eights share only 6 of 8 members. Per-capita is a *different question*, not a trick that flips a map. The figure also documents the caveat the other two cannot: a rate map of a *helpline* measures service reach as much as need, and the map says so itself: its darkest county is **McKean, at 2 requests against 39,904 residents** - 0.05 per 1,000, where the next-lowest county (Elk) is 12.3, some 245 times higher. Nobody believes McKean has no hardship; it sits at the edge of *this* call center's intake, so the figure is measuring who dials 2-1-1, not who needs it. A per-capita map inherits whatever its numerator was actually counting. Nothing is stored in the CSV but the requests themselves: `--geojson auto` fetches the county boundaries **and** canonicalizes the county *names* to Census GEOIDs, and `--denominator census@2024` fetches ACS 5-year total population (`B01003`), so there is no FIPS column and no committed boundary file. The upstream feed's `zip_code` is dropped rather than mis-tagged (a second geo concept would add a competing choropleth candidate), and the dataset's older 2020-2023 resource is deliberately not concatenated: 5.4% of its rows carry a pipe-delimited `Allegheny County|Westmoreland County` multi-county value, and every way of resolving those changes the map. Built by the new `examples/viz/gen_wpa_211_requests.py`, which asserts its own row count, county count and date span so a silently revised upstream feed fails instead of quietly re-cutting the figure the caption describes. diff --git a/tests/test_sqlp.rs b/tests/test_sqlp.rs index 1884115eed..0620d7799e 100644 --- a/tests/test_sqlp.rs +++ b/tests/test_sqlp.rs @@ -5969,6 +5969,12 @@ fn sqlp_window_order_by_multiple_keys_limitation() { // Mixing directions or NULLS placement across window keys is rejected // outright. If a future polars fixes any of this, these assertions fail and // the tiebreak workaround becomes available. + // + // Reported upstream as pola-rs/polars#29390. Root cause: + // Expr::over_with_options collapses several ORDER BY keys into a single + // as_struct(...), and a struct holding a null field is not itself null, so + // SortOptions::nulls_last has nothing to act on. `descending` survives the + // same path, which is why only the null placement is wrong. let wrk = Workdir::new("sqlp_window_order_by_multiple_keys_limitation"); window_nulls_fixture(&wrk); From 49631062771827be4f1608784c8084956e0877b4 Mon Sep 17 00:00:00 2001 From: Joel Natividad <1980690+jqnatividad@users.noreply.github.com> Date: Sat, 19 Sep 2026 13:54:16 -0400 Subject: [PATCH 4/4] docs: name the two ISO-based date-part functions instead of "Both" The bullet lists 11 date-part functions and then said "Both are ISO-based", which reads as if it refers to the whole list. It means WEEK and DAYOFWEEK, so say so. Co-Authored-By: Claude Opus 5 (1M context) --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index fb8cfc9619..c4f6fece50 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,7 +9,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Added - **`sqlp`: new SQL surface from the Polars bump to `9d5804d`** (upstream `py-1.44.0` -> the pinned rev). None of this needed qsv changes beyond the bump, but none of it was covered by tests either, so `tests/test_sqlp.rs` now exercises each one: - **`GROUP BY GROUPING SETS (...)` / `ROLLUP(...)` / `CUBE(...)`**, plus **`GROUPING(k, ...)`** and **`GROUPING_ID(k, ...)`** to tell a subtotal row's NULL marker apart from a NULL in the data ([#29278](https://github.com/pola-rs/polars/pull/29278)). Both pack one bit per argument, MSB first, so argument order matters. Pair them with `--wnull-value` - a subtotal's NULL keys are otherwise written as empty fields. - - **the bare date-part functions `YEAR`, `QUARTER`, `MONTH`, `WEEK`, `DAY`/`DAYOFMONTH`, `DAYOFWEEK`, `DAYOFYEAR`, `HOUR`, `MINUTE`, `SECOND`** ([#29269](https://github.com/pola-rs/polars/pull/29269)). Both are ISO-based, which is the part worth knowing: `WEEK('2024-12-31')` is **1** (ISO week 1 of 2025), not 53, and `DAYOFWEEK` counts Monday as 1. + - **the bare date-part functions `YEAR`, `QUARTER`, `MONTH`, `WEEK`, `DAY`/`DAYOFMONTH`, `DAYOFWEEK`, `DAYOFYEAR`, `HOUR`, `MINUTE`, `SECOND`** ([#29269](https://github.com/pola-rs/polars/pull/29269)). `WEEK` and `DAYOFWEEK` are ISO-based, which is the part worth knowing: `WEEK('2024-12-31')` is **1** (ISO week 1 of 2025), not 53, and `DAYOFWEEK` counts Monday as 1. - **`DATE '...'` / `TIMESTAMP '...'` typed literals** ([#29007](https://github.com/pola-rs/polars/pull/29007)), and **date +/- integer arithmetic** shifting by whole days with the integer on either side, from a literal or a column, plus **Decimal comparisons in non-equi joins** ([#29156](https://github.com/pola-rs/polars/pull/29156)). - **parenthesized `JOIN ... ON (...)` constraints** ([#28967](https://github.com/pola-rs/polars/pull/28967)); relation aliases declared *inside* those parens now resolve in the outer SELECT, since bare parens around a join do not open a new scope ([#29158](https://github.com/pola-rs/polars/pull/29158)). - **`OVER` on multi-argument aggregates** - `CORR`, `COVAR_POP`, `COVAR_SAMP`, `QUANTILE_CONT`, `QUANTILE_DISC`, `STRING_AGG` ([#29160](https://github.com/pola-rs/polars/pull/29160)) - and **`NULLS FIRST`/`NULLS LAST` inside a window's own `ORDER BY`** ([#29159](https://github.com/pola-rs/polars/pull/29159)) - though only when the window has a **single** sort key: with two or more keys the `NULLS` clause is silently ignored and the nulls follow the sort direction instead, which we reported upstream as [#29390](https://github.com/pola-rs/polars/issues/29390).