Segfault on large multi-column Iceberg upserts
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 64/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Tranquilla
- Stack tecnologico
- python
- Ambito
- data-engineering, databases
Direzione di ricerca
Inizia da pyiceberg.table.upsert_util.create_match_filter ed esegui iceberg_upsert_segfault_repro.py per riprodurre il crash dell’upsert su più colonne, usando pyiceberg-stacktrace.txt per confermare il percorso di canonicalizzazione. Raggruppa le tuple di chiavi in un numero minore di disgiunzioni preservando le corrispondenze esatte, quindi verifica che il caso sintetico a bassa cardinalità non vada più in segfault e valuta la limitazione relativa all’alta cardinalità.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Apache Iceberg version
0.11.0 (latest release)
Please describe the bug 🐞
When upserting into an Iceberg table, PyIceberg first scans the target table to
find which existing rows match the source rows' key columns. It builds that
"matching" predicate in pyiceberg.table.upsert_util.create_match_filter:
-
For a single join column it emits one flat
In(col, [v1, v2, ...]).
PyArrow lowers this to a singleis_incompute node, no matter how many
values it contains — so single-column upserts of huge tables are fine. -
For a multi-column key it instead emits one disjunct per distinct key
tuple::Or(And(c1 == v1, c2 == w1), And(c1 == v2, c2 == w2), ...) # ONE disjunct PER ROW
PyIceberg builds that Or as a balanced tree, so the Python side copes.
But when the expression is handed to PyArrow's dataset scanner as a filter, the
C++ expression engine canonicalises it: Dataset::GetFragments calls
SimplifyWithGuarantee → Canonicalize, which flattens the associative
or_kleene chain and then recurses over it. With tens of thousands of
disjuncts that recursion overflows the C++ call stack and the process
segfaults (SIGSEGV) — typically after several minutes of work, with a
backtrace full of arrow::compute::Canonicalize / ModifyExpression
frames.
Reference: https://github.com/apache/iceberg-python/issues/3272
Note that apache/iceberg-python#3448 addresses a different upsert segfault (a
per-batch Acero re-filter in _task_to_record_batches, mostly observed on
Apple Silicon). It does not touch the GetFragments canonicalisation path
exercised here, so it does not help with this crash.
The fix
Produce a predicate that matches exactly the same rows, but with far fewer
disjuncts. Group the key tuples and emit a single In over whichever column
collapses to the fewest distinct "prefix" combinations (choosing that column
makes the result independent of the caller's column ordering)::
Or(And(c1 == v1, c2 IN [w, x, y]),
And(c1 == v2, c2 IN [z]),
...) # one disjunct per distinct PREFIX
The disjunct count drops from "number of rows" to "number of distinct prefix
values". In the synthetic data below there are 50 000 unique ids spread over
just 50 group values, so the predicate shrinks from 50 000 disjuncts to 50 —
shallow enough that PyArrow's canonicaliser no longer overflows.
Caveat
This helps whenever at least one key column is low-cardinality (or, equivalently,
one column is near-unique and can be folded into the In). A genuinely
high-cardinality composite key — where every column is near-unique and all of
them are needed to identify a row — still produces roughly one disjunct per row
even after grouping, and can still overflow. For that pathological case the
only robust option is to upsert in smaller batches.
iceberg_upsert_segfault_repro.py
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- Lingua principale
- Python
- Stelle
- 1.1k
- Fork
- 589
- Merge medio
- 2g 2h
- PR unite (30g)
- 70
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di apache/iceberg-python
-
kind:bug
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
apache/iceberg-python#4006 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
apache/iceberg-python#3996 ·
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validation Apertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
apache/iceberg-python#3979 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
apache/iceberg-python#3885 ·
-
[Bug] PyArrowFileIO fails to propagate s3.ssl.ca-cert to pyarrow.fs.S3FileSystem tls_ca_file_path Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
apache/iceberg-python#3866 · 1 commento ·
Tutte le issue di apache/iceberg-python
Issue simili
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
canonical/paas-charm#368 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
tech debt
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
addition to tracking list Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
StevenBlack/hosts#3256 ·
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
qualcomm/qai-appbuilder#275 ·