Upsert with 1M rows extremely slow due to `create_match_filter` and `txn.delete()` performance
維護者通常 1 天內回覆
評估
- 難度
- 5/5
- 預估耗時
- 一週以上
- 新手友好度
- 35/100
- Issue 類型
- 缺陷
- 描述清晰度
- 需要釐清
- 活躍度
- 冷清
- 技術堆疊
- python
研究方向
從報告中描述的 create_match_filter 和 txn.delete() 進入點開始,然後根據提供的計時結果重現 1M 列 upsert 基準測試。查看相關 issue #2159、#2138 和 #2943,以了解現有內容。完成標準是取得經測量的改善,或提供一種有文件記錄且受支援的方法,以避免報告中的整個資料表刪除成本。
由索引模型根據 Issue 內容生成。
描述
Apache Iceberg version
0.11.0
Please describe the bug 🐞
CC @goutamvenkat-anyscale @koenvo @Fokko
Hello! We are implementing distributed writes from Ray Data to Iceberg. As part of upserts, we:
- Write data files in parallel across Ray workers (each worker writes its share of Parquet files directly to storage and returns
DataFilemetadata + the upsert key columns back to the driver) - On the driver, concatenate all upsert keys collected from workers, call
create_match_filterto build a delete predicate, then calltxn.delete()followed by an append to commit
Upserting 1M rows (383 MiB) into an Iceberg table takes ~17.5 minutes, almost entirely in the delete step:
create_match_filter (1M keys → In filter): 10.26s
txn.delete(): 1054.35s
append + commit: 1.14s
─────────────────────────────────────────────────────
Total upsert commit: 1065.75s
PyIceberg version 0.11.0
This matches what's reported in #2159 and #2138.
The bottlenecks are:
create_match_filter— constructs a PythonBooleanExpressionnode per row, which is expensive at 1M+ keystxn.delete()— evaluates the resulting giantInexpression against the table's data files with no partition pruning, effectively doing a full table scan
We have a few questions:
- Merge-on-read upserts — is this on the roadmap, and if so, roughly when? MoR would let us avoid the expensive delete + rewrite cycle entirely for large upserts.
- Optimizing
create_match_filterortxn.delete()— is there a recommended way to speed these up today? For example, batching theInfilter, or passing a partition-level hint to constrain the file scan? - Partition-aware deletes — if the upsert key columns overlap with partition columns, is there a supported way to restrict
txn.delete()to only the relevant partitions, rather than scanning the full table?
Related
- #2159 — Upserting large table extremely slow
- #2138 — Upsertion memory usage grows exponentially as table size grows
- #2943 — Optimize upsert performance for large datasets
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- 主要語言
- Python
- 星號
- 1.2k
- 分支
- 618
- 平均合併
- 1 天 10 小時
- 30 天內合併 PR
- 71
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 有 Pull Request 範本
- 沒有貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
apache/iceberg-python 的其他 Issue
-
難度 2/5 1-3 小時 新手友好度 70/100
apache/iceberg-python#4093 ·
維護者通常 1 天內回覆
-
Replace `__slots__ = (field1,field2,...)` with `slots=True`可能已有人在做 @med9110 於 2 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 68/100
apache/iceberg-python#4086 · 3 則留言 ·
維護者通常 1 天內回覆
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's response可能已有人在做 @Soumo-git-hub 於 3 天前認領。 未關閉kind:bug
難度 2/5 1-3 小時 新手友好度 84/100
apache/iceberg-python#4073 · 1 則留言 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 70/100
apache/iceberg-python#4010 · 3 則留言 · 1 個 reaction ·
維護者通常 1 天內回覆
-
to_bytes silently rescales a Decimal with a negative scale可能已有人在做 @Rodrigo-Palma 於 22 天前認領。 未關閉
難度 2/5 1-3 小時 新手友好度 78/100
apache/iceberg-python#3996 ·
維護者通常 1 天內回覆
查看 apache/iceberg-python 的全部 Issue
相似的 Issue
-
area:space-accuracy good first issue track:data
難度 2/5 1-3 小時 新手友好度 85/100
Sara-Managed-Projects/space-radar#904 ·
維護者通常 1 天內回覆