When pyiceberg loads Iceberg tables containing large JSON data into memory, memory usage explodes in pyarrow
Les mainteneurs répondent en général sous 1 jour
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Accessibilité débutants
- 48/100
Piste de recherche
Commencez par les points d’entrée table.scan(...).to_arrow() et to_arrow_batch_reader(), puis suivez la manière dont ils invoquent pyarrow. Comparez cela avec pyarrow.parquet.read_table et son option read_dictionary. Le travail est terminé lorsque les appelants peuvent configurer l’encodage par dictionnaire pour des colonnes sélectionnées sans modifier le comportement par défaut, l’option étant transmise lors des lectures.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Apache Iceberg version
0.11.0 (latest release)
Please describe the bug 🐞
We have a table with a column containing large JSON strings (basically think of large pydantic validation results). The resulting datafile in Iceberg itself is a roughly 6MB (GZip compressed) Parquet file, but when querying it, the memory consumption goes to ~4GB. (Just for a few seconds, but long enough to cause out-of-memory issues on some systems.)
The issue is that by default pyarrow loads the strings per row into memory, which blows up the memory. If we download the datafile and open it directly via pyarrow this behaviour can be reproduced.
There are 2 workarounds in pyarrow: 1. Don't load the problematic column (given that is possible in your use case) and 2. switch to dictionary-encoding for set column (example snippet below).
from pyarrow.parquet import read_table
table = read_table("datafile.parquet", read_dictionary=["problematic_column"])
Issue in pyiceberg:
Regardless if you use table.scan(...).to_arrow() or table.scan(...).to_arrow_batch_reader(), pyiceberg has afaik currently no option to specify the dictionary encoding for certain tables, hence pyarrow uses the default encoding and the memory usage explodes.
The to_arrow_batch_reader does not help here either, because -as per my understanding- in the batch reader of pyiceberg each batch represents an individual datafile. Hence, if there is one problematic 6MB datafile, it makes no difference if you use the batch reader or not. I also have the impression that when you iterate over the reader, pyarrow has already loaded the parquet file in a separate thread and this is where the memory explosion actually happens.
So the current only workaround in pyiceberg is option 1: Don't load the problematic column by specifying the selected_fields:
from pyiceberg.catalog import load_catalog
catalog = load_catalog("default")
table = catalog.load_table("your_table")
reader = table.scan(selected_fields=("all", "other", "columns")).to_arrow_batch_reader()
for batch in reader:
...
Expected behaviour:
There should be an option somewhere, e.g. in the data_scan to specify for which columns dictionary encoding should be used. This option should be forwarded to pyarrow internally somehow, so that pyarrow uses less memory.
Remark:
I would not change the default behaviour. It would be just good to have the option to configure the encoding in pyarrow when needed.
This issue is a follow up for https://github.com/apache/iceberg-python/issues/1205
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- Langage dominant
- Python
- Étoiles
- 1.1k
- Forks
- 606
- Merge moyen
- 1 j 18 h
- PR mergées (30 j)
- 87
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Propose un modèle de pull request
- Aucun guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de apache/iceberg-python
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's responseOuvertekind:bug
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100
apache/iceberg-python#4073 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
apache/iceberg-python#4010 · 3 commentaires · 1 réaction ·
Les mainteneurs répondent en général sous 1 jour
-
to_bytes silently rescales a Decimal with a negative scalePeut-être pris @Rodrigo-Palma l’a pris il y a 17 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3996 ·
Les mainteneurs répondent en général sous 1 jour
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationPeut-être pris @ghoshp83 l’a pris il y a 17 jours. Ouvertebug
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
apache/iceberg-python#3979 ·
Les mainteneurs répondent en général sous 1 jour
-
FsspecFileIO: `_adls` mutates shared properties, so a second storage account gets the first account's filesystemPeut-être pris @krishnakaanchan-png l’a pris il y a 34 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3885 ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de apache/iceberg-python
Issues similaires
-
Broken links found in docsOuvertedocs pydanty:is-working
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
pydantic/pydantic-ai#9800 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
-
[Bug]: With --api-server-count > 1, gauges such as vllm:num_requests_running have no samples until the first requestPeut-être pris @roy6n23 l’a pris aujourd’hui. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
vllm-project/vllm#59988 · 2 commentaires ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
pymc-devs/pymc-examples#897 ·
-
Action calls retired claude-3-5-haiku-20241022, generating failing API requests for every userOuverte
Difficulté 1/5 Moins d'une heure Accessibilité débutants 91/100