[SPARK-48308][Core] Unify getting data schema without partition columns in FileSourceStrategy #46619

johanl-db · 2024-05-16T12:47:21Z

What changes were proposed in this pull request?

Compute the schema of the data without partition columns only once in FileSourceStrategy.

Why are the changes needed?

In FileSourceStrategy, the schema of the data excluding partition columns is computed 2 times in a slightly different way, using an AttributeSet (partitionSet) and using the attributes directly (partitionColumns)
These don't have the exact same semantics, AttributeSet will only use expression ids for comparison while comparing with the actual attributes will use the name, type, nullability and metadata. We want to use the former here.

Does this PR introduce any user-facing change?

No

How was this patch tested?

Existing tests

Was this patch authored or co-authored using generative AI tooling?

No

cloud-fan · 2024-05-16T14:37:34Z

thanks, merging to master!

…columns in FileSourceStrategy Compute the schema of the data without partition columns only once in FileSourceStrategy. In FileSourceStrategy, the schema of the data excluding partition columns is computed 2 times in a slightly different way, using an AttributeSet (`partitionSet`) and using the attributes directly (`partitionColumns`) These don't have the exact same semantics, AttributeSet will only use expression ids for comparison while comparing with the actual attributes will use the name, type, nullability and metadata. We want to use the former here. No Existing tests No Closes apache#46619 from johanl-db/reuse-schema-without-partition-columns. Authored-by: Johan Lasperas <[email protected]> Signed-off-by: Wenchen Fan <[email protected]>

Reuse schema without partition columns in FileSourceStrategy

7387068

github-actions bot added the SQL label May 16, 2024

cloud-fan approved these changes May 16, 2024

View reviewed changes

cloud-fan closed this in 57948c8 May 16, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

[SPARK-48308][Core] Unify getting data schema without partition columns in FileSourceStrategy #46619

[SPARK-48308][Core] Unify getting data schema without partition columns in FileSourceStrategy #46619

Uh oh!

johanl-db commented May 16, 2024

Uh oh!

cloud-fan commented May 16, 2024

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

[SPARK-48308][Core] Unify getting data schema without partition columns in FileSourceStrategy #46619

[SPARK-48308][Core] Unify getting data schema without partition columns in FileSourceStrategy #46619

Uh oh!

Conversation

johanl-db commented May 16, 2024

What changes were proposed in this pull request?

Why are the changes needed?

Does this PR introduce any user-facing change?

How was this patch tested?

Was this patch authored or co-authored using generative AI tooling?

Uh oh!

cloud-fan commented May 16, 2024

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants