[Feature] TrinoSplitManager calls dropStats() before scan planning, preventing manifest-level file pruning
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 75/100
- Issue type
- Feature
- Clarity
- Clearly specified
- Activity status
- Quiet
- Tech stack
- java
- Domain
- data-engineering
Research direction
Start in src/main/java/org/apache/paimon/trino/TrinoSplitManager.java at the readBuilder.dropStats().newScan().plan().splits() call. Review how dropStats() affects planning and confirm that predicate evaluation can use manifest statistics before stats are removed from the resulting splits. Done means metadata.stats-mode enables file-level pruning while split serialization still omits unnecessary statistics.
Written by the indexing model from the issue text.
Description
Search before asking
- I searched in the issues and found nothing similar.
Motivation
metadata.stats-mode rendered useless for Trino
Users who configure metadata.stats-mode = 'full' or metadata.stats-mode = 'truncate(16)' receive no benefit in Trino. Statistics are written to manifests correctly but discarded before they can be used for pruning.
Solution
Summary
TrinoSplitManager calls .dropStats() before newScan().plan(), which strips manifest-level column statistics from DataFileMeta entries before predicate evaluation can use them. As a result, Trino cannot use metadata.stats-mode min/max statistics for file-level skipping — all files become splits regardless of predicate filters.
.dropStats() was added deliberately in commit 5144ad9 when upgrading from Paimon 0.8.0 to 1.0-SNAPSHOT, likely to reduce split serialisation overhead or fix a serialisation error. The fix is not to remove it, but to move it to after predicate evaluation — so statistics are used for pruning during plan() and then stripped before splits are sent to workers.
Root Cause
File: src/main/java/org/apache/paimon/trino/TrinoSplitManager.java, line 86
// Current — dropStats() called BEFORE scan planning; stats unavailable for predicate evaluation
List<Split> splits = readBuilder.dropStats().newScan().plan().splits();
.dropStats() flags the scan to zero out SimpleStats (min/max/null-count per column per file) on each DataFileMeta entry. Because it is called before newScan().plan(), the predicate filter wired via readBuilder.withFilter() has no statistics to evaluate against during plan(). Every file passes the statistics check (vacuously, since stats are empty) and becomes a split.
History
.dropStats() was introduced in Paimon core in #4506 (November 2024) as an explicit optimisation to reduce the size of split objects sent to workers — splits carry DataFileMeta entries, which include column statistics that are not needed by workers after planning. Spark added dropStats() in #5093 in a secondary path. When paimon-trino upgraded to 1.0-SNAPSHOT in 5144ad9, .dropStats() was added to TrinoSplitManager — the commit message ("Update Paimon core to 1.0-SNAPSHOT / fix") gives no further explanation, but the intent is the same serialisation optimisation.
The problem is placement: .dropStats() must be called after plan() completes, not before. Calling it before plan() eliminates the serialisation overhead but also eliminates all statistics-based file pruning.
What dropStats() does
ReadBuilder.dropStats() sets a flag that causes AbstractFileStoreScan to call DataFileMeta.copyWithoutStats() on each entry — replacing column statistics with EMPTY_STATS before returning results. The flag is evaluated during plan(). If set before plan(), stats are zeroed before predicate evaluation. If set after plan() (on the returned splits), stats are zeroed only for serialisation, after pruning has already occurred.
Anything else?
No response
Are you willing to submit a PR?
- I'm willing to submit a PR!
- Dominant language
- Java
- Stars
- 3.4k
- Forks
- 1.4k
- Avg merge
- 1d 14h
- Merged PRs (30d)
- 468
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/paimon
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
[Bug] [Hive] IndexOutOfBoundsException when converting an unavailable dynamic BETWEEN predicate Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
infinispan/infinispan#18150 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
untriaged
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
opensearch-project/k-NN#3597 ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100