-
Notifications
You must be signed in to change notification settings - Fork 31
DataFusion 52 release post #135
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from 2 commits
Commits
Show all changes
30 commits
Select commit
Hold shift + click to select a range
68abcfb
Initial draft (coded with codex)
alamb 1686945
updates
alamb 63b8e12
Merge remote-tracking branch 'apache/main' into site/datafusion_52
alamb 3c7dd6a
updates
alamb 2cc1f59
Update sql planning
alamb ccc5d42
Apply suggestions from code review
alamb 81954c5
Updates
alamb 781cd62
acknowledgments
alamb 790b658
update
alamb d63a4de
updates
alamb d38d99f
update
alamb b47c50d
typos
alamb 1f5b91e
refine
alamb 34cea38
Merge branch 'site/datafusion_52' of https://github.com/apache/datafu…
alamb 2823de5
clean
alamb 1c9dadf
Update content/blog/2026-01-08-datafusion-52.0.0.md
alamb 615affd
Clarify RelationPlanner syntax
alamb 8329e6a
Merge branch 'site/datafusion_52' of https://github.com/apache/datafu…
alamb 1345bfb
remove extra section
alamb 63b571c
Update content/blog/2026-01-08-datafusion-52.0.0.md
alamb 7984011
Refine wording
alamb 66a9dcb
reflow
alamb e9308d4
Metadata cache is general
alamb 4e24b1f
Add section on Min/max dynamic filters
alamb d62dfab
Update content/blog/2026-01-08-datafusion-52.0.0.md
alamb 1a77f71
whitespace
alamb 859d40d
Merge branch 'site/datafusion_52' of https://github.com/apache/datafu…
alamb 37b12dd
Improve sort pushdown description
alamb 91f71f3
capitalizaton
alamb 040b85c
wordsmith
alamb File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Some comments aren't visible on the classic Files Changed page.
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,348 @@ | ||
| --- | ||
| layout: post | ||
| title: Apache DataFusion 52.0.0 Released | ||
| date: 2026-01-08 | ||
| author: pmc | ||
| categories: [release] | ||
| --- | ||
|
|
||
| <!-- | ||
| {% comment %} | ||
| Licensed to the Apache Software Foundation (ASF) under one or more | ||
| contributor license agreements. See the NOTICE file distributed with | ||
| this work for additional information regarding copyright ownership. | ||
| The ASF licenses this file to you under the Apache License, Version 2.0 | ||
| (the "License"); you may not use this file except in compliance with | ||
| the License. You may obtain a copy of the License at | ||
|
|
||
| http://www.apache.org/licenses/LICENSE-2.0 | ||
|
|
||
| Unless required by applicable law or agreed to in writing, software | ||
| distributed under the License is distributed on an "AS IS" BASIS, | ||
| WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| See the License for the specific language governing permissions and | ||
| limitations under the License. | ||
| {% endcomment %} | ||
| --> | ||
|
|
||
| [TOC] | ||
|
|
||
| ## Introduction | ||
|
|
||
| We are proud to announce the release of [DataFusion 52.0.0]. This post highlights | ||
| some of the major improvements since [DataFusion 51.0.0]. The complete list of | ||
| changes is available in the [changelog]. Thanks to the [120 contributors] for | ||
| making this release possible. | ||
|
|
||
| TODO: confirm the release date for 52.0.0 and update the front matter if needed. | ||
|
alamb marked this conversation as resolved.
Outdated
|
||
|
|
||
| [DataFusion 52.0.0]: https://crates.io/crates/datafusion/52.0.0 | ||
| [DataFusion 51.0.0]: https://datafusion.apache.org/blog/2025/11/25/datafusion-51.0.0/ | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. TODO |
||
| [changelog]: https://github.com/apache/datafusion/blob/branch-52/dev/changelog/52.0.0.md | ||
| [120 contributors]: https://github.com/apache/datafusion/blob/branch-52/dev/changelog/52.0.0.md#credits | ||
|
|
||
| ## Performance Improvements 🚀 | ||
|
|
||
| We continue to make significant performance improvements in DataFusion, both in | ||
| the core engine and in the Parquet reader. This release includes faster `CASE` | ||
| expressions, better hash performance for string types, and continued string | ||
| function optimizations. | ||
|
mbutrovich marked this conversation as resolved.
Outdated
|
||
|
|
||
| ### Performance Chart (TODO) | ||
|
|
||
| TODO: add the 52.0.0 performance chart and update the caption. | ||
|
|
||
| <img | ||
| src="/blog/images/datafusion-52.0.0/performance_over_time_clickbench.png" | ||
| width="100%" | ||
| class="img-responsive" | ||
| alt="Performance over time" | ||
| /> | ||
|
|
||
| **Figure 1**: TODO: update caption for 52.0.0 benchmarking results. | ||
|
|
||
| ## Major Features ✨ | ||
|
|
||
| ### Arrow IPC Stream file support | ||
|
|
||
| DataFusion can now read Arrow IPC stream files ([#18457]). This expands | ||
| interoperability with systems that emit Arrow streams directly, making it | ||
| simpler to ingest Arrow-native data without conversion. | ||
|
|
||
| Example (TODO: confirm exact syntax for IPC stream format selection): | ||
|
|
||
| ```sql | ||
| -- TODO: confirm whether the format name is `arrow`, `ipc_stream`, or implicit. | ||
| CREATE EXTERNAL TABLE ipc_events | ||
| STORED AS ARROW | ||
| LOCATION 's3://bucket/events.arrow'; | ||
| ``` | ||
|
|
||
| Related PRs: [#18457] | ||
|
|
||
| [#18457]: https://github.com/apache/datafusion/pull/18457 | ||
|
|
||
| ### Faster `CASE` expression evaluation | ||
|
|
||
| DataFusion 52 completes major work from the CASE performance epic ([#18075]). | ||
|
alamb marked this conversation as resolved.
Outdated
|
||
| Lookup-table based evaluation avoids repeated expression evaluation and reduces | ||
| branching overhead, accelerating common ETL patterns. | ||
|
|
||
| Example: | ||
|
|
||
| ```sql | ||
| SELECT | ||
| CASE | ||
| WHEN status IN ('NEW', 'READY', 'STAGED') THEN 'PENDING' | ||
| WHEN status IN ('DONE', 'COMPLETE') THEN 'FINISHED' | ||
| ELSE 'OTHER' | ||
| END AS status_bucket, | ||
| count(*) | ||
| FROM jobs | ||
| GROUP BY 1; | ||
| ``` | ||
|
|
||
| Related PRs: [#18183] | ||
|
|
||
| [#18075]: https://github.com/apache/datafusion/issues/18075 | ||
| [#18183]: https://github.com/apache/datafusion/pull/18183 | ||
|
|
||
| ### Extensible SQL planning with relation planner extensions | ||
|
|
||
| DataFusion now supports relation planner extensions for custom SQL syntax and | ||
| planning logic ([#17824], [#17843]). This lets downstream projects inject their | ||
| own planning behavior without forking the SQL planner, which is critical for | ||
| dialect extensions and custom table references. | ||
|
|
||
| Diagram: | ||
|
|
||
| ``` | ||
| SQL text | ||
| | (custom relation planner extension) | ||
| v | ||
| Logical plan | ||
| | (DataFusion optimizers) | ||
| v | ||
| Physical plan | ||
| ``` | ||
|
|
||
| TODO: include a short Rust snippet showing how to register a relation planner | ||
| extension once the final API example is confirmed. | ||
|
|
||
| Related PRs: [#17843] | ||
|
|
||
| [#17824]: https://github.com/apache/datafusion/issues/17824 | ||
| [#17843]: https://github.com/apache/datafusion/pull/17843 | ||
|
|
||
| ### ListingTable object store usage improvements | ||
|
|
||
| ListingTable improvements continue to reduce object store I/O and planning | ||
| latency for partitioned datasets ([#17214]). DataFusion now normalizes partition | ||
| and flat listings, enables a memory-bound list-files cache by default, and | ||
| makes the cache prefix-aware for partition pruning. | ||
|
|
||
| Diagram: | ||
|
|
||
| ``` | ||
| Object store LIST | ||
| | (normalized listing + cache) | ||
| v | ||
| Partitioned files | ||
| | (planner) | ||
| v | ||
| Execution plan | ||
| ``` | ||
|
|
||
| Related PRs: [#18146], [#18855], [#19366], [#19298], [#18971] | ||
|
|
||
| [#17214]: https://github.com/apache/datafusion/issues/17214 | ||
| [#18146]: https://github.com/apache/datafusion/pull/18146 | ||
| [#18855]: https://github.com/apache/datafusion/pull/18855 | ||
| [#19366]: https://github.com/apache/datafusion/pull/19366 | ||
| [#19298]: https://github.com/apache/datafusion/pull/19298 | ||
| [#18971]: https://github.com/apache/datafusion/pull/18971 | ||
|
|
||
| ### Statistics cache improvements | ||
|
|
||
| The statistics cache has been improved to make pruning and planning more | ||
| reliable in repeated workloads ([#19051]). DataFusion now exposes a | ||
| `statistics_cache` function and improves cache memory behavior for listing | ||
| workflows, making it easier to diagnose cache contents and reduce repeated I/O. | ||
|
|
||
| Example (TODO: confirm the function signature and output schema): | ||
|
|
||
| ```sql | ||
| -- TODO: confirm the function name and arguments. | ||
| SELECT * FROM statistics_cache('my_table'); | ||
| ``` | ||
|
|
||
| Related PRs: [#19054], [#18855], [#18971] | ||
|
|
||
| [#19051]: https://github.com/apache/datafusion/issues/19051 | ||
| [#19054]: https://github.com/apache/datafusion/pull/19054 | ||
|
|
||
| ### Pushdown expression evaluation via PhysicalExprAdapter | ||
|
|
||
| DataFusion now pushes down expression evaluation into TableProviders using the | ||
| PhysicalExprAdapter, replacing the older SchemaAdapter approach ([#14993], | ||
| [#16800]). This enables richer pushdown (expressions and projections) and | ||
| improves consistency between logical and physical planning. | ||
|
|
||
| Diagram: | ||
|
|
||
| ``` | ||
| SQL filter/projection | ||
| | (PhysicalExprAdapter) | ||
| v | ||
| TableProvider pushdown | ||
| | (scan) | ||
| v | ||
| Reduced data | ||
| ``` | ||
|
|
||
| Related PRs: [#18998], [#19345] | ||
|
|
||
| [#14993]: https://github.com/apache/datafusion/issues/14993 | ||
| [#16800]: https://github.com/apache/datafusion/issues/16800 | ||
| [#18998]: https://github.com/apache/datafusion/pull/18998 | ||
| [#19345]: https://github.com/apache/datafusion/pull/19345 | ||
|
|
||
| ### Hash join build-side pushdown | ||
|
|
||
| DataFusion can now push down build-side hash tables from HashJoinExec into scans | ||
| ([#17171]). When the build side is small, DataFusion converts the hash table to | ||
| an `IN` list or hash lookup that can be evaluated during scans, reducing the | ||
| join input size early. | ||
|
|
||
| Example: | ||
|
|
||
| ```sql | ||
| SELECT * | ||
| FROM orders o | ||
| JOIN small_dim d | ||
| ON o.dim_id = d.id; | ||
| ``` | ||
|
|
||
| TODO: include a physical plan snippet that shows the pushdown filter once a | ||
| canonical example is selected. | ||
|
|
||
| Related PRs: [#18393] | ||
|
|
||
| [#17171]: https://github.com/apache/datafusion/issues/17171 | ||
| [#18393]: https://github.com/apache/datafusion/pull/18393 | ||
|
|
||
| ### Sort pushdown to sources | ||
|
|
||
| DataFusion now supports sort pushdown into data sources, allowing scans to | ||
| return sorted data or leverage reversed row groups when possible ([#10433], | ||
| [#19064]). This reduces memory pressure and can eliminate explicit sort stages | ||
| for partitioned or pre-sorted data. | ||
|
|
||
| Example: | ||
|
|
||
| ```sql | ||
| SELECT * | ||
| FROM parquet_table | ||
| ORDER BY event_time DESC; | ||
| ``` | ||
|
|
||
| Related PRs: [#19064] | ||
|
|
||
| [#10433]: https://github.com/apache/datafusion/issues/10433 | ||
| [#19064]: https://github.com/apache/datafusion/pull/19064 | ||
|
|
||
| ### DELETE/UPDATE hooks in TableProvider | ||
|
|
||
| TableProvider now includes DELETE and UPDATE hooks, with MemTable providing the | ||
| first implementation ([#19142]). This is an important step toward fully | ||
| featured DML support and enables downstream storage engines to plug in their | ||
| own mutation logic. | ||
|
|
||
| Example: | ||
|
|
||
| ```sql | ||
| DELETE FROM mem_table WHERE status = 'obsolete'; | ||
| ``` | ||
|
|
||
| Related PRs: [#19142] | ||
|
|
||
| [#19142]: https://github.com/apache/datafusion/pull/19142 | ||
|
|
||
| ### CoalesceBatchesExec removal and integrated batch coalescing | ||
|
|
||
| DataFusion continues the work from the CoalesceBatchesExec epic ([#18779]). The | ||
| standalone `CoalesceBatchesExec` operator existed to ensure batches were large | ||
| enough for vectorized execution, and it was inserted after filter-like | ||
| operators such as `FilterExec`, `HashJoinExec`, and `RepartitionExec`. However, | ||
| it also blocked other optimizations (like pushing limits through joins) and | ||
| made optimizer rules more complex. This release integrates coalescing into the | ||
| operators themselves and relies on Arrow's coalesce kernels, reducing plan | ||
| complexity while keeping batch sizes efficient. | ||
|
|
||
| Diagram: | ||
|
|
||
| ``` | ||
| Before: | ||
| Scan -> CoalesceBatches -> Filter -> CoalesceBatches -> Join | ||
|
|
||
| After: | ||
| Scan -> Filter (coalesce inline) -> Join (coalesce inline) | ||
| ``` | ||
|
|
||
| Related PRs: [#18540], [#18604], [#18630], [#18972], [#19002], [#19342], [#19239] | ||
| Thanks to [Tim-53], [Dandandan], [jizezhang], and [feniljain] for implementing | ||
| this feature. | ||
|
|
||
| [#18779]: https://github.com/apache/datafusion/issues/18779 | ||
| [#18540]: https://github.com/apache/datafusion/pull/18540 | ||
| [#18604]: https://github.com/apache/datafusion/pull/18604 | ||
| [#18630]: https://github.com/apache/datafusion/pull/18630 | ||
| [#18972]: https://github.com/apache/datafusion/pull/18972 | ||
| [#19002]: https://github.com/apache/datafusion/pull/19002 | ||
| [#19342]: https://github.com/apache/datafusion/pull/19342 | ||
| [#19239]: https://github.com/apache/datafusion/pull/19239 | ||
| [Tim-53]: https://github.com/Tim-53 | ||
| [Dandandan]: https://github.com/Dandandan | ||
| [jizezhang]: https://github.com/jizezhang | ||
| [feniljain]: https://github.com/feniljain | ||
|
|
||
| ## Upgrade Guide and Changelog | ||
|
|
||
| Upgrading to 52.0.0 should be straightforward for most users. Please review the | ||
| [Upgrade Guide] | ||
| for details on breaking changes and code snippets to help with the transition. | ||
| For a comprehensive list of all changes, please refer to the [changelog]. | ||
|
|
||
| ## About DataFusion | ||
|
|
||
| [Apache DataFusion] is an extensible query engine, written in [Rust], that uses | ||
| [Apache Arrow] as its in-memory format. DataFusion is used by developers to | ||
| create new, fast, data-centric systems such as databases, dataframe libraries, | ||
| and machine learning and streaming applications. While [DataFusion's primary | ||
| design goal] is to accelerate the creation of other data-centric systems, it | ||
| provides a reasonable experience directly out of the box as a [dataframe | ||
| library], [Python library], and [command-line SQL tool]. | ||
|
|
||
| [apache datafusion]: https://datafusion.apache.org/ | ||
| [rust]: https://www.rust-lang.org/ | ||
| [apache arrow]: https://arrow.apache.org | ||
| [DataFusion's primary design goal]: https://datafusion.apache.org/user-guide/introduction.html#project-goals | ||
| [dataframe library]: https://datafusion.apache.org/user-guide/dataframe.html | ||
| [python library]: https://datafusion.apache.org/python/ | ||
| [command-line SQL tool]: https://datafusion.apache.org/user-guide/cli/ | ||
| [Upgrade Guide]: https://datafusion.apache.org/library-user-guide/upgrading.html | ||
|
|
||
| ## How to Get Involved | ||
|
|
||
| DataFusion is not a project built or driven by a single person, company, or | ||
| foundation. Rather, our community of users and contributors works together to | ||
| build a shared technology that none of us could have built alone. | ||
|
|
||
| If you are interested in joining us, we would love to have you. You can try out | ||
| DataFusion on some of your own data and projects and let us know how it goes, | ||
| contribute suggestions, documentation, bug reports, or a PR with documentation, | ||
| tests, or code. A list of open issues suitable for beginners is [here], and you | ||
| can find out how to reach us on the [communication doc]. | ||
|
|
||
| [here]: https://github.com/apache/arrow-datafusion/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22 | ||
| [communication doc]: https://datafusion.apache.org/contributor-guide/communication.html | ||
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
According to:
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Thanks -- I think in the past we have dated the blog posts based on when the post was released rather than when the software was 🤔