When pyiceberg loads Iceberg tables containing large JSON data into memory, memory usage explodes in pyarrow
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Facilidade para iniciantes
- 48/100
Direção de pesquisa
Comece pelos pontos de entrada table.scan(...).to_arrow() e to_arrow_batch_reader() e, em seguida, rastreie como eles invocam pyarrow. Compare isso com pyarrow.parquet.read_table e sua opção read_dictionary. Considera-se concluído quando os chamadores puderem configurar a codificação de dicionário para colunas selecionadas sem alterar o comportamento padrão, com a opção sendo encaminhada durante as leituras.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Apache Iceberg version
0.11.0 (latest release)
Please describe the bug 🐞
We have a table with a column containing large JSON strings (basically think of large pydantic validation results). The resulting datafile in Iceberg itself is a roughly 6MB (GZip compressed) Parquet file, but when querying it, the memory consumption goes to ~4GB. (Just for a few seconds, but long enough to cause out-of-memory issues on some systems.)
The issue is that by default pyarrow loads the strings per row into memory, which blows up the memory. If we download the datafile and open it directly via pyarrow this behaviour can be reproduced.
There are 2 workarounds in pyarrow: 1. Don't load the problematic column (given that is possible in your use case) and 2. switch to dictionary-encoding for set column (example snippet below).
from pyarrow.parquet import read_table
table = read_table("datafile.parquet", read_dictionary=["problematic_column"])
Issue in pyiceberg:
Regardless if you use table.scan(...).to_arrow() or table.scan(...).to_arrow_batch_reader(), pyiceberg has afaik currently no option to specify the dictionary encoding for certain tables, hence pyarrow uses the default encoding and the memory usage explodes.
The to_arrow_batch_reader does not help here either, because -as per my understanding- in the batch reader of pyiceberg each batch represents an individual datafile. Hence, if there is one problematic 6MB datafile, it makes no difference if you use the batch reader or not. I also have the impression that when you iterate over the reader, pyarrow has already loaded the parquet file in a separate thread and this is where the memory explosion actually happens.
So the current only workaround in pyiceberg is option 1: Don't load the problematic column by specifying the selected_fields:
from pyiceberg.catalog import load_catalog
catalog = load_catalog("default")
table = catalog.load_table("your_table")
reader = table.scan(selected_fields=("all", "other", "columns")).to_arrow_batch_reader()
for batch in reader:
...
Expected behaviour:
There should be an option somewhere, e.g. in the data_scan to specify for which columns dictionary encoding should be used. This option should be forwarded to pyarrow internally somehow, so that pyarrow uses less memory.
Remark:
I would not change the default behaviour. It would be just good to have the option to configure the encoding in pyarrow when needed.
This issue is a follow up for https://github.com/apache/iceberg-python/issues/1205
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- Linguagem predominante
- Python
- Estrelas
- 1.2k
- Forks
- 618
- Merge médio
- 1d 18h
- PRs com merge (30d)
- 70
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Tem um modelo de pull request
- Sem guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de apache/iceberg-python
-
PyArrowFileIO: every small S3 write is a 3-request multipart upload; expose allow_delayed_openAberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 70/100
apache/iceberg-python#4093 ·
Mantenedores costumam responder em até 1 dia
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's responseTalvez já em andamento @Soumo-git-hub assumiu há 3 dias. Abertakind:bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 84/100
apache/iceberg-python#4073 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 70/100
apache/iceberg-python#4010 · 3 comentários · 1 reação ·
Mantenedores costumam responder em até 1 dia
-
to_bytes silently rescales a Decimal with a negative scaleTalvez já em andamento @Rodrigo-Palma assumiu há 23 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
apache/iceberg-python#3996 ·
Mantenedores costumam responder em até 1 dia
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationTalvez já em andamento @ghoshp83 assumiu há 23 dias. Abertabug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
apache/iceberg-python#3979 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de apache/iceberg-python
Issues semelhantes
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 60/100
521xueweihan/HelloGitHub#3924 ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 67/100
wilbowes/EchoMuse#869 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 85/100
-
Claiming namespace `jft63`Abertanamespace operations
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 72/100
EclipseFdn/open-vsx.org#14043 ·
Mantenedores costumam responder em até 1 dia
-
test: TestServeUntilStale races the server's close against the client's sendall (BrokenPipeError under load)Talvez já em andamento @evoludigit assumiu hoje. Aberta
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 89/100
Mantenedores costumam responder em até 1 dia