Feature: Add metadata-only replace API to Table for REPLACE snapshot operations
还没有人认领这个 Issue。
评估
调研方向
从 pyiceberg/table/update/snapshot.py 和现有的 Table 与 Transaction snapshot API 开始,然后将拟议行为与 Java 的 RewriteFiles 接口进行比较。检查 tests/table/test_snapshots.py,尤其是 test_invalid_operation(),并为文件交换、序列号和 operation=REPLACE 添加有针对性的覆盖测试。当仅元数据替换能够使用 Iterable[DataFile] 输入以原子方式工作并避免 Parquet 序列化时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
Feature Request / Improvement
Description
This issue proposes implementing a metadata-only replace API in PyIceberg, enabling orchestrators to submit a set of DataFiles to delete and a set of DataFiles to append in a single atomic transaction.
This functionality is critical for maintenance operations such as data compaction (the "small files" problem), ensuring the logical state of the table remains unaltered while physical data layout is optimized.
Background
In a current PR (#3124, part of #1092), PyIceberg's replace semantics are tightly coupled with PyArrow dataframes (def replace(self, df: pa.Table)). This approach introduces several architectural flaws:
- Coupling Physical Serialization with Metadata: It forces a
.parquetwrite serialization hook directly into the snapshot commit transaction, increasing the risk of schema degradation and blocking network topologies. - Missing
Operation.REPLACE: The current system uses primitives that log asAPPENDorOVERWRITE, muddying the table history and complicating snapshot expiry/maintenance. - Java Inconsistency: This severely drifts from Java Iceberg's native
org.apache.iceberg.RewriteFilesspecification, which strictly isolates the builder into accepting purelyDataFilepointers.
Proposed Solution
To fix this and achieve logical equivalence, we must implement an exact port of Java's RewriteFiles builder API into PyIceberg's native _SnapshotProducer engine.
-
Introduce
_RewriteFilesSnapshot Producer:
Add a new_RewriteFilesclass that specifically targets replacing existing files. This class will implement:_deleted_entries(): To find the existing target files and re-emit them asDELETEDentries, defensively keeping their ancestralsequence_numbers completely intact for time travel compatibility._existing_manifests(): To scavenge unchanged manifests natively, skipping deep rewrites and only mutating manifests impacted by the deleted files.
-
Builder Hook Implementation:
ImplementUpdateSnapshot().replace()which configures the transaction withOperation.REPLACE. -
Expose Shorthands on Table & Transaction:
AddreplaceAPIs on bothTableandTransactiontakingIterable[DataFile]arguments to elegantly wrap the snapshot mutation:def replace( self, files_to_delete: Iterable[DataFile], files_to_add: Iterable[DataFile], snapshot_properties: dict[str, str] = EMPTY_DICT, branch: str | None = MAIN_BRANCH, ) -> None: ...
Notable canges
replace()API implemented on bothTableandTransactionusingIterable[DataFile].- PyArrow
.parquetwrite logic decoupled from the metadata transaction. _RewriteFilescorrectly copies ancestralsequence_numberpointers forDELETEDandEXISTINGmanifest entries.- Snapshots committed via the
replace()hook possess a Summary containingoperation=Operation.REPLACE. - Unit tests pass simulating data file swaps and summary verifications.
Related Java API
Inspired heavily by Java's builder interface: https://github.com/apache/iceberg/blob/main/api/src/main/java/org/apache/iceberg/RewriteFiles.java
AI Disclosure
AI was used to help understand the code base and draft code changes. All code changes have been thoroughly reviewed, ensuring that the code changes are in line with a broader understanding of the codebase.
- Worth deeper review after AI-assistance:
- The
test_invalid_operation()intests/table/test_snapshots.pypreviously usedOperation.REPLACEas a value to test invalid operations, but with this changeOperation.REPLACEbecomes valid. In place I just put a dummy Operation. - The
_RewriteFilesinpyiceberg/table/update/snapshot.pyoverrides the_deleted_entriesand_existing_manifestsfunctions. I sought to test this thoroughly that it was done correctly. I am thinking it's possible to improve the test suite to make this more rigorous. I am open to suggestions on how that could be done.
- 主要语言
- Python
- 星标
- 1.1k
- 派生
- 589
- 平均合并
- 2 天 2 小时
- 30 天内合并 PR
- 70
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/iceberg-python 的其他 Issue
-
kind:bug
难度 1/5 1 小时以内 新手友好度 92/100
apache/iceberg-python#4006 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
apache/iceberg-python#3996 ·
-
bug
难度 2/5 1-3 小时 新手友好度 72/100
apache/iceberg-python#3979 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
apache/iceberg-python#3885 ·
-
[Bug] PyArrowFileIO fails to propagate s3.ssl.ca-cert to pyarrow.fs.S3FileSystem tls_ca_file_path 未关闭
难度 2/5 1-3 小时 新手友好度 76/100
apache/iceberg-python#3866 · 1 条评论 ·
查看 apache/iceberg-python 的全部 Issue
相似的 Issue
-
enhancement
难度 2/5 1-3 小时 新手友好度 70/100
canonical/paas-charm#368 · 1 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
-
tech debt
难度 2/5 1-3 小时 新手友好度 75/100
-
难度 1/5 1 小时以内 新手友好度 90/100
StevenBlack/hosts#3256 ·
-
难度 1/5 1 小时以内 新手友好度 90/100
qualcomm/qai-appbuilder#275 ·