ArrowScan to_table fails if the data is mixed between dict-encoded strings and plain strings.
Les mainteneurs répondent en général sous 1 jour
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 3/5
- Temps estimé
- 1-2 jours
- Accessibilité débutants
- 55/100
Piste de recherche
Commencez par la reproduction minimale et inspectez ArrowScan.to_table ainsi que le point d’entrée to_record_batches présenté dans l’issue. Vérifiez comment sont combinés les batches mêlant des chaînes simples et des valeurs encodées par dictionnaire. Le travail est terminé lorsque la reproduction s’achève sans ArrowTypeError et renvoie une table Arrow avec la colonne de chaînes attendue.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Apache Iceberg version
0.11.0 (latest release)
Please describe the bug 🐞
We have recently updated our functions to call the pyiceberg table.append() function with dict encoded arrow tables. Now we have in our iceberg tables mixed data from before this change, (where our data still is stored as string) and after the change, where the data is stored as dict-encoded strings.
If we now call to_arrow() of a DataScan class, on this table we get this error:
pyarrow.lib.ArrowTypeError: Unable to merge: Field col has incompatible types: string vs dictionary<values=string, indices=int32, ordered=0>
Here is a minimal example that reproduces this error:
from pyiceberg.io.pyarrow import ArrowScan
from pyiceberg.table import ALWAYS_TRUE
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField
from pyiceberg.types import StringType
import pyarrow as pa
def create_scan_with_mixed_dict_encode_not_encode() -> ArrowScan:
schema = Schema(
NestedField(field_id=1, name="col", field_type=StringType(), required=False)
)
class FakeTableMetadata:
def schema(self) -> Schema:
return schema
scan = ArrowScan(table_metadata=FakeTableMetadata(),
io=object(),
projected_schema=schema,
row_filter=ALWAYS_TRUE)
def _batches_for_repro(self, _tasks):
str_values = pa.array(["a"], type=pa.string())
yield pa.record_batch([str_values], names=["col"])
yield pa.record_batch([str_values.dictionary_encode()], names=["col"])
ArrowScan.to_record_batches = _batches_for_repro
return scan
if __name__ == "__main__":
scan = create_scan_with_mixed_dict_encode_not_encode()
arrow_table = ArrowScan.to_table(scan, tasks=[])
I am happy to provide a bugfix PR, but I need a small guidance on the best approach.
One idea is to cast each batch in to_table to the arrow_schema. The more performant way is to check for each batch, if the schema is different. If they are different, then find the dict_encoded col and only cast that one to string.
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- Langage dominant
- Python
- Étoiles
- 1.1k
- Forks
- 589
- Merge moyen
- 2 j 11 h
- PR mergées (30 j)
- 67
Préparer son environnement
Nous n'avons pas encore vérifié les fichiers d'installation de ce projet. Commencez par son README, et consultez notre guide de la première contribution pour les étapes générales.
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de apache/iceberg-python
-
asf-allowlist-check fails on every PR: the pinned setup-uv ref expired from the ASF allowlistOuvertekind:bug
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
apache/iceberg-python#4026 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
apache/iceberg-python#4010 · 3 commentaires · 1 réaction ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3996 ·
Les mainteneurs répondent en général sous 1 jour
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationOuvertebug
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
apache/iceberg-python#3979 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3885 ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de apache/iceberg-python
Issues similaires
-
bug status/needs-triage
Difficulté 2/5 1-3 heures Accessibilité débutants 86/100
prowler-cloud/prowler#12887 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
-
area: desktop platform: macos priority: p3 status: ready type: enhancement
Difficulté 1/5 Moins d'une heure Accessibilité débutants 92/100
use-agent-os/agent-os#3484 ·
Les mainteneurs répondent en général sous 2 jours
-
bug
Difficulté 2/5 1-3 heures Accessibilité débutants 86/100
open-telemetry/opentelemetry-python-contrib#5113 · 2 commentaires · 2 réactions ·
Les mainteneurs répondent en général sous 1 jour
-
external
Difficulté 2/5 1-3 heures Accessibilité débutants 68/100
langchain-ai/docs#6255 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
Les mainteneurs répondent en général sous 1 jour