[8.x] First step optimizing tsdb doc values codec merging. by martijnvg · Pull Request #126827 · elastic/elasticsearch

martijnvg · 2025-04-15T09:22:33Z

Backporting #125403 to the 8.x branch.

The doc values codec iterates a few times over the doc value instance that needs to be written to disk. In case when merging and index sorting is enabled, this is much more expensive, as each time the doc values instance is iterated a merge sorting is performed (in order to get the doc ids of new segment in order of index sorting).

There are several reasons why the doc value instance is iterated multiple times:

To compute stats (num values, number of docs with value) required for writing values to disk.
To write bitset that indicate which documents have a value. (indexed disi, jump table)
To write the actual values to disk.
To write the addresses to disk (in case docs have multiple values)

This applies for numeric doc values, but also for the ordinals of sorted (set) doc values.

This PR addresses solving the first reason why doc value instance needs to be iterated. This is done only when in case of merging and when the segments to be merged with are also of type es87 doc values, codec version is the same and there are no deletes. Note this optimized merged is behind a feature flag for now.

Backporting elastic#125403 to the 8.x branch. The doc values codec iterates a few times over the doc value instance that needs to be written to disk. In case when merging and index sorting is enabled, this is much more expensive, as each time the doc values instance is iterated a merge sorting is performed (in order to get the doc ids of new segment in order of index sorting). There are several reasons why the doc value instance is iterated multiple times: * To compute stats (num values, number of docs with value) required for writing values to disk. * To write bitset that indicate which documents have a value. (indexed disi, jump table) * To write the actual values to disk. * To write the addresses to disk (in case docs have multiple values) This applies for numeric doc values, but also for the ordinals of sorted (set) doc values. This PR addresses solving the first reason why doc value instance needs to be iterated. This is done only when in case of merging and when the segments to be merged with are also of type es87 doc values, codec version is the same and there are no deletes. Note this optimized merged is behind a feature flag for now.

The compatibleWithOptimizedMerge() method doesn't handle codec readers that are wrapped by our source pruning filter codec reader. This change addresses that. Failing to detect this means that the optimized merge will not kick in.

martijnvg added backport v8.19.0 labels Apr 15, 2025

martijnvg added 4 commits April 15, 2025 11:46

fixed compile errors in benchmark

01ead9f

Fix DocValuesConsumerUtil (elastic#126836)

c4e6478

The compatibleWithOptimizedMerge() method doesn't handle codec readers that are wrapped by our source pruning filter codec reader. This change addresses that. Failing to detect this means that the optimized merge will not kick in.

Merge remote-tracking branch 'es/8.x' into backport_125403

481fe32

Merge remote-tracking branch 'es/8.x' into backport_125403

1da9ed2

martijnvg added the auto-merge-without-approval Automatically merge pull request when CI checks pass (NB doesn't wait for reviews!) label Apr 16, 2025

Merge remote-tracking branch 'es/8.x' into backport_125403

ecc2a41

elasticsearchmachine merged commit e9dcb50 into elastic:8.x Apr 17, 2025
15 checks passed

martijnvg deleted the backport_125403 branch April 17, 2025 07:23

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[8.x] First step optimizing tsdb doc values codec merging.#126827

[8.x] First step optimizing tsdb doc values codec merging.#126827
elasticsearchmachine merged 6 commits intoelastic:8.xfrom
martijnvg:backport_125403

martijnvg commented Apr 15, 2025

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Comments

Conversation

martijnvg commented Apr 15, 2025

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Comments