[HUDI-5253] HoodieMergeOnReadTableInputFormat could have duplicate records issue if it contains delta files while still splittable#7264
Merged
codope merged 3 commits intoapache:masterfrom Nov 29, 2022
Conversation
danny0405
reviewed
Nov 22, 2022
hudi-hadoop-mr/src/main/java/org/apache/hudi/hadoop/realtime/HoodieRealtimePath.java
Show resolved
Hide resolved
Contributor
Author
|
@hudi-bot run azure |
1 similar comment
Contributor
Author
|
@hudi-bot run azure |
Contributor
Author
|
@danny0405 @xushiyan could you please take a look? |
danny0405
reviewed
Nov 29, 2022
| HoodieLogFile logFile = new HoodieLogFile(fs.getFileStatus(new Path(logPath))); | ||
| rtPath = new HoodieRealtimePath(new Path("foo"), "bar", basePath.toString(), Collections.singletonList(logFile), "000", false, Option.empty()); | ||
| assertFalse(new HoodieMergeOnReadTableInputFormat().isSplitable(fs, rtPath), "Path for bootstrap should not be splitable."); | ||
| } |
Contributor
There was a problem hiding this comment.
Path for bootstrap should not be splitable
Is the error message right ?
danny0405
reviewed
Nov 29, 2022
| void pathNotSplitableIfContainsDeltaFiles() throws IOException { | ||
| URI basePath = Files.createTempFile(tempDir, "target", ".parquet").toUri(); | ||
| HoodieRealtimePath rtPath = new HoodieRealtimePath(new Path("foo"), "bar", basePath.toString(), Collections.emptyList(), "000", false, Option.empty()); | ||
| assertTrue(new HoodieMergeOnReadTableInputFormat().isSplitable(fs, rtPath)); |
Contributor
There was a problem hiding this comment.
Supplement the error message.
Collaborator
satishkotha
pushed a commit
that referenced
this pull request
Dec 13, 2022
…cords issue if it contains delta files while still splittable (#7264)
alexeykudinkin
pushed a commit
to onehouseinc/hudi
that referenced
this pull request
Dec 14, 2022
…cords issue if it contains delta files while still splittable (apache#7264)
alexeykudinkin
pushed a commit
to onehouseinc/hudi
that referenced
this pull request
Dec 14, 2022
…cords issue if it contains delta files while still splittable (apache#7264)
alexeykudinkin
pushed a commit
to onehouseinc/hudi
that referenced
this pull request
Dec 14, 2022
…cords issue if it contains delta files while still splittable (apache#7264)
alexeykudinkin
pushed a commit
to onehouseinc/hudi
that referenced
this pull request
Dec 14, 2022
…cords issue if it contains delta files while still splittable (apache#7264)
alexeykudinkin
pushed a commit
to onehouseinc/hudi
that referenced
this pull request
Dec 14, 2022
…cords issue if it contains delta files while still splittable (apache#7264)
fengjian428
pushed a commit
to fengjian428/hudi
that referenced
this pull request
Apr 5, 2023
…cords issue if it contains delta files while still splittable (apache#7264)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Change Logs
If
HoodieRealtimePathcontains delta files, it cannot splittableImpact
We can also find that sometimes could throw {{IllegalStateException}} duplicates key error when we run the CI.
We can easily to reproduce this in
org.apache.hudi.testutils.HoodieMergeOnReadTestUtils#getRecordsUsingInputFormatto allow it create more splitsRisk level (write none, low medium or high below)
low
Documentation Update
Describe any necessary documentation update if there is any new feature, config, or user-facing change
ticket number here and follow the instruction to make
changes to the website.
Contributor's checklist