feat: Add native fixed-array dot - #28504
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #28504 +/- ##
==========================================
+ Coverage 81.53% 81.56% +0.02%
==========================================
Files 1883 1884 +1
Lines 266376 266521 +145
Branches 3224 3227 +3
==========================================
+ Hits 217188 217384 +196
+ Misses 48355 48305 -50
+ Partials 833 832 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
c9347ef to
c28fe48
Compare
|
Possible follow-up, separate from this PR, is to fuse sibling df.select(
score_a=pl.col("embedding").arr.dot(query_a),
score_b=pl.col("embedding").arr.dot(query_b),
)Today each expression traverses embedding independently. Planner could read each embedding array once, compute both scores in one multi-query kernel, and still return two separately named columns. Expressions that cannot be fused would continue using existing Standalone Apple M4 experiment found this faster from two query vectors onward, but it has not been validated inside Polars and floating-point tolerance is unresolved. I don't want to mess with baseline. |
07fad23 to
1f2765f
Compare
1f2765f to
16dce81
Compare
|
Addressed raw-query-vector API mismatch. |
d3df4a5 to
3e79306
Compare
|
@orlp , @ritchie46 , whenever you have time, this one is ready. I don't want to grow scope, it's reasonable with limitations and follow-up work. Happy to answer your questions. |
|
Verified current implementation explicitly on both produced identical results |
564905d to
9d28983
Compare
e7dac5b to
a284fad
Compare
|
It’s ready for review, but it’s currently blocked by new, unrelated CI bugs. I have to convert it back to a draft due to my two-PR limit so I can prioritize CI fix. Feel free to review this PR anyway. @wence- If you want to see additional measurements, let me know which. |
Allow Python lists, 1D NumPy arrays, and literal vectors to inherit the left Array dtype, matching embedding-scoring workflows without weakening strict dtype checks between columns.
Document null and floating-point behavior, and cover all-missing rows, special values, cancellation, fragmented inputs, and wide all-valid arrays.
9d845e8 to
9c62320
Compare
Closes #17456 and benefits embedding pipelines doing similarity scoring or candidate reranking.
Related #24652 is tracked separately by #28614: vertical fixed-array aggregation Expr.sum() reduces Array rows into one Array, while arr.dot reduces each row into one scalar.
Issue example uses
Expr.dot, which reduces columns to one scalar. This PR addsarr.dotbecause requested operation is row-wise over each fixed-size array.API distinction:
Existing equivalent materializes a
rows × widthmultiplication result before reducing:Native kernel fuses multiplication and reduction, so it does not materialize rows × width product. Fragmented inputs still rechunk, so peak memory can include contiguous copies of lhs and rhs in addition to scalar output.
Attached benchmark in comments.
Scope:
Float32orFloat64inner dtypes;Limitations:
Update (2026-07-30):
arr.dotnow accepts Python sequences and one-dimensional NumPy arrays directly. These raw query vectors are cast after selector or wildcard expansion, when each left expression has concrete Array dtype. Explicit expressions and Series operands retain their dtype and must match the left operand.AI disclosure:
OpenAI Codex drafted early benchmark code and plots with my revision, fixed my grammar and phrasing.