IN predicates > 200 values disable file pruning
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 58/100
Hướng nghiên cứu
Start in _InclusiveMetricsEvaluationVisitor.visit_in and inspect the related use in table/update/validate.py; reproduce the 200-versus-201 benchmark described in the issue. Compare file-pruning behavior for large IN predicates and check whether the same evaluator affects conflict detection. Done means large predicates retain safe bounds-based pruning without changing the existing precise path below 200 values.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Apache Iceberg version
main (development)
Please describe the bug 🐞
While profiling the delete path from #3129, I noticed that plan_files() stops using file bounds as soon as an IN predicate contains more than 200 values.
In _InclusiveMetricsEvaluationVisitor.visit_in, we return ROWS_MIGHT_MATCH before checking the file bounds:
keys=200 planned_files=1 /20 rewritten=1
keys=201 planned_files=20 /20 rewritten=1
keys=1000 planned_files=20 /20 rewritten=1
This was on an unpartitioned table with 20 files of 10k rows each, disjoint key ranges, with all deleted keys belonging to the first file.
So the 201st value does not make the predicate significantly more expensive — it disables pruning and makes every file a candidate.
I also see the same effect when scaling the table: with the same layout, the 200 → 201 transition added about 0.15s on 5 files vs 3.60s on 80 files (median of 7 runs).
The limit comes from #1588 / #1672. The original concern was that evaluating large IN predicates could cost more than the pruning saves. However, the bounds check can potentially be reduced to a single min() / max() computation per predicate instead of scanning all literals for every file.
I would keep the current precise evaluation below the 200-value limit and only restore bounds-based pruning above it, where we currently don't prune at all.
One more thing: the same evaluator is used by conflict detection in table/update/validate.py, so this may also affect false-positive conflicts for large IN predicates.
Questions
Is disabling all pruning above 200 values intentional?
Would a min/max bounds check above the limit be acceptable?
Should the Python change be mirrored in Java?
I haven't implemented the fix yet; I'd rather confirm the intended behavior first.
Benchmark: local SQLite catalog, Python 3.10, pyarrow 25.0.1, pyiceberg 0d58407, median of 7 runs.
I used an AI assistant to help run the benchmarks and inspect the code path; the reproduction and measurements are mine.
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- Ngôn ngữ chính
- Python
- Star
- 1.1k
- Fork
- 589
- Merge trung bình
- 1 ngày 20 giờ
- Pull request đã merge (30 ngày)
- 68
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/iceberg-python
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
apache/iceberg-python#4010 · 1 reaction ·
-
kind:bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
apache/iceberg-python#4006 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/iceberg-python#3996 ·
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validation Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
apache/iceberg-python#3979 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/iceberg-python#3885 ·
Tất cả issue của apache/iceberg-python
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
xinnan-tech/xiaozhi-fde-talk#263 ·
-
rules
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
huggingface/Repo2RLEnv#163 · 1 bình luận ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 95/100
huggingface/sentence-transformers#4074 ·
-
comp/dashboard invalid P3
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
NousResearch/hermes-agent#121143 ·