Bug: dynamic_partition_overwrite silently skips spec-0 manifests after partition spec evolution
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Accessibilité débutants
- 70/100
Piste de recherche
Commencez dans table/init.py, au niveau de Table.dynamic_partition_overwrite, puis suivez _build_partition_projection et la modification du manifest-pruning de #3011. Reproduisez le snapshot mixte spec-0/spec-1 décrit dans l’issue et comparez-le au correctif associé précédent de #1108. Le travail est terminé lorsqu’un test de régression confirme que l’écrasement supprime les lignes correspondantes des deux specs historiques sans supprimer les données sans rapport.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Summary
dynamic_partition_overwrite produces incorrect results when a table has undergone
partition spec evolution. Manifests written under older specs are silently skipped
by the manifest pruning logic introduced in #3011, leaving stale data files that
should have been deleted.
Root cause
In Table.dynamic_partition_overwrite (table/__init__.py), the delete predicate
is built using only the current partition spec:
delete_filter = self._build_partition_predicate(
partition_records=partitions_to_overwrite,
spec=self.table_metadata.spec(), # always current spec
schema=self.table_metadata.schema()
)
A snapshot with mixed partition_spec_ids (spec-0 and spec-1 manifests) passes
this single predicate to _DeleteFiles. The manifest evaluator in _build_partition_projection
uses inclusive_projection(schema, spec) — when projecting a spec-1 predicate
(e.g. category=A AND region=us) through spec-0 (which only has category), the
region reference has no corresponding partition field, causing the evaluator to
incorrectly skip spec-0 manifests entirely.
Reproduction
import tempfile, pyarrow as pa
from pyiceberg.catalog import load_catalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, StringType, LongType
from pyiceberg.partitioning import PartitionSpec, PartitionField
from pyiceberg.transforms import IdentityTransform
schema = Schema(
NestedField(1, "category", StringType(), required=False),
NestedField(2, "region", StringType(), required=False),
NestedField(3, "value", LongType(), required=False),
)
spec_v0 = PartitionSpec(
PartitionField(source_id=1, field_id=1000, transform=IdentityTransform(), name="category")
)
with tempfile.TemporaryDirectory() as warehouse:
catalog = load_catalog("test", **{"type": "sql", "uri": f"sqlite:///{warehouse}/catalog.db", "warehouse": f"file://{warehouse}"})
catalog.create_namespace("default")
table = catalog.create_table("default.test", schema=schema, partition_spec=spec_v0)
# Write under spec 0
table.append(pa.table({"category": ["A","A","B"], "region": [None,None,None], "value": [1,2,10]}))
# Evolve spec
with table.update_spec() as u:
u.add_field("region", IdentityTransform(), "region")
table = catalog.load_table("default.test")
# Write under spec 1
table.append(pa.table({"category": ["A","B"], "region": ["us","us"], "value": [100,200]}))
# Overwrite category=A — should delete ALL A rows (both specs)
table.dynamic_partition_overwrite(
pa.table({"category": ["A"], "region": ["us"], "value": [999]})
)
result = table.scan().to_arrow().to_pydict()
a_values = [v for c,v in zip(result["category"], result["value"]) if c == "A"]
print(a_values) # BUG: prints [1, 2, 100, 999] — stale rows from spec-0 not deleted
# EXPECTED: [999]
Fix
Build the delete predicate per historical spec present in the snapshot, projecting
the new data files' partition values into each spec's coordinate space before evaluating.
PR with fix and regression tests to follow.
Related
- #3011 (introduced the manifest pruning optimization)
- #1108 (prior related fix by @Fokko for spec evolution in manifest rewriting)
- Langage dominant
- Python
- Étoiles
- 1.1k
- Forks
- 589
- Merge moyen
- 1 j 20 h
- PR mergées (30 j)
- 68
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de apache/iceberg-python
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
apache/iceberg-python#4010 · 1 réaction ·
-
kind:bug
Difficulté 1/5 Moins d'une heure Accessibilité débutants 92/100
apache/iceberg-python#4006 ·
-
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3996 ·
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validation Ouvertebug
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
apache/iceberg-python#3979 ·
-
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3885 ·
Toutes les issues de apache/iceberg-python
Issues similaires
-
bug
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
stephrobert/dsoxlab#238 ·
-
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
-
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
sublimehq/package_control#1780 ·
-
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
nwg-piotr/nwg-displays#145 ·