Support merge-on-read deletes on v2 tables
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 38/100
Hướng nghiên cứu
Bắt đầu từ nhánh merge-on-read của Transaction.delete() và kiểm tra POSITIONAL_DELETE_SCHEMA, delete-file index hiện có và v2 manifest writer. Truy vết cách manifests và snapshots được tạo, sau đó xác minh rằng position resolution đọc các data-file rows khớp theo physical order. Hoàn thành khi v2 opt-in deletes tạo exact-bound position-delete files trong DELETES manifests, trong khi behavior copy-on-write mặc định của v1 và v3 vẫn không thay đổi.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Feature Request / Improvement
Transaction.delete() is copy-on-write: deleting rows rewrites every data file that holds a matched row, so it costs ~O(table) per operation and grows with the table. Add v2 merge-on-read position-delete writes so a delete instead writes a small position-delete file, costing O(rows deleted), with the data files left untouched. v2 position deletes are readable by the existing query engines today, so this is broadly adoptable.
Even as v3 merge-on-read (deletion vectors) support lands, v2 position-delete writes remain valuable, because format version is a per-table property and the two encodings are mutually exclusive: a v2 table can never use a deletion vector. v3 write support therefore does nothing for the large existing v2 fleet, whose tables stay v2 until deliberately upgraded and often must stay v2 to remain readable by every engine that consumes them (v3 deletion vectors require newer readers). For those tables, position deletes are the only efficient delete path, and they are also the format Java/Spark Iceberg already writes, so this also closes a read/write asymmetry in PyIceberg. Deletion vectors are the right default for new v3 tables; v2 position deletes serve the installed base.
What already exists (no work needed)
- The read path already applies position deletes at scan time (the delete-file index / position-delete application).
POSITIONAL_DELETE_SCHEMA(file_path,pos),DataFileContent.POSITION_DELETES,ManifestContent.DELETES, and the full snapshot / manifest-list / commit machinery.Transaction.delete()already recognizeswrite.delete.mode=merge-on-read. It currently just warns "not yet supported" and falls back to copy-on-write. That fallback is the hook point.
The gap to close
- Widen
posinPOSITIONAL_DELETE_SCHEMAfrominttolong. The Iceberg spec typesposaslong; keeping itintmismatches what other engines write/read and overflows on files with more than 2^31 rows. The read path treatsposas a generic index, so widening is safe. - A DELETES manifest writer. The v2 manifest writer hardcodes manifest content to
data; a variant is needed that writes contentdeletes. - A snapshot producer that adds position-delete files (writes a DELETES manifest for them, grouped by partition spec, and carries existing manifests forward), committing an overwrite snapshot.
- A position-delete file writer that records exact
file_pathlower/upper bounds. The read-side delete-file index pins a delete to a single data file only when itsfile_pathbound is exact; the defaulttruncate(16)write-metrics mode would truncate that bound and mis-route (or over-apply) the delete. - Predicate to positions resolution. For each data file a delete predicate touches, read it in physical order and collect the ordinal row positions that match, then write one position-delete file per touched data file.
- Wire 1 through 5 into the merge-on-read branch of
Transaction.delete(), gated behindwrite.delete.mode=merge-on-readon v2 tables; copy-on-write stays the default so nothing changes unless a user opts in. v1 (cannot store delete manifests) and v3 (deletion vectors, out of scope) keep the copy-on-write fallback. A size-based heuristic (rewrite tiny files instead of writing delete files) can come later.
- Ngôn ngữ chính
- Python
- Star
- 1.1k
- Fork
- 589
- Merge trung bình
- 1 ngày 20 giờ
- Pull request đã merge (30 ngày)
- 68
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/iceberg-python
-
kind:bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
apache/iceberg-python#4006 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/iceberg-python#3996 ·
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validation Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
apache/iceberg-python#3979 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/iceberg-python#3885 ·
-
[Bug] PyArrowFileIO fails to propagate s3.ssl.ca-cert to pyarrow.fs.S3FileSystem tls_ca_file_path Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
apache/iceberg-python#3866 · 1 bình luận ·
Tất cả issue của apache/iceberg-python
Issue tương tự
-
bug confirmed issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
open-webui/open-webui#30750 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
OpenwaterHealth/openmotion-bloodflow-app#604 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
good first issue
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100