Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

PyArrowFileIO: every small S3 write is a 3-request multipart upload; expose allow_delayed_open

Abierto Apto para principiantes
#4,093 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
2/5
Tiempo estimado
1-3 horas
Aptitud para principiantes
70/100
Tipo de issue
Nueva funcionalidad
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
aws, python
Área
cloud, data

Línea de trabajo

Empieza en _initialize_s3_fs en PyArrowFileIO, donde el sistema de archivos S3 se construye a partir de un conjunto fijo de propiedades. Revisa cómo se leen y se pasan las propiedades s3.* existentes. Añade una propiedad s3.allow-delayed-open, pasa allow_delayed_open solo cuando el pyarrow instalado sea la 21 o posterior, y usa true como valor por defecto, tal como propone el issue. Está terminado cuando las escrituras pequeñas salen como un único PutObject y la opción no aparece en pyarrow más antiguo.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Apache Iceberg version

0.12.0 (also on main)

Please describe the bug 🐞

PyArrowFileIO writes every S3 object as a multipart upload, however small it is. pyarrow's S3FileSystem.open_output_stream starts a multipart upload as soon as the stream opens (apache/arrow#51029). Each metadata.json, manifest list, manifest, version-hint.text and small data file therefore costs three requests: CreateMultipartUpload, UploadPart and CompleteMultipartUpload. All three are billed as writes, even for a 1-byte object.

pyarrow's S3FileSystem has an option for this, allow_delayed_open, since pyarrow 21. With it set, a stream that closes before reaching a part's size is sent as a single PutObject, and a larger one is still a multipart upload. _initialize_s3_fs builds the filesystem from a fixed set of properties, though, so there's no way to set the option through FileIO properties.

The cost is real for a table that commits often. Each commit writes several small metadata objects, so it makes about three times the write requests it needs.

Repro

Run against any S3-compatible endpoint (this was run against a local rustfs):

import pyarrow.fs as fs
from pyiceberg.io.pyarrow import PyArrowFileIO

fs.initialize_s3(fs.S3LogLevel.Debug)  # logs each request

io = PyArrowFileIO({
    "s3.endpoint": "http://127.0.0.1:9000",
    "s3.access-key-id": "...",
    "s3.secret-access-key": "...",
    "s3.region": "us-east-1",
})
with io.new_output("s3://bucket/version-hint.text").create(overwrite=True) as f:
    f.write(b"1")

The debug log shows three requests for the one byte:

POST /bucket/version-hint.text?uploads
PUT  /bucket/version-hint.text?partNumber=1&uploadId=...
POST /bucket/version-hint.text?uploadId=...

When the same write goes through an S3FileSystem built with allow_delayed_open=True, the log shows a single PUT /bucket/version-hint.text.

Proposal

Pass allow_delayed_open in _initialize_s3_fs, controlled by a FileIO property such as s3.allow-delayed-open. I'd suggest defaulting it to true, since it only changes how small objects are uploaded. pyiceberg supports pyarrow 18 and up, so the option would be passed only on pyarrow 21 or later.

The workaround today is to subclass PyArrowFileIO, override _initialize_s3_fs, and rebuild the filesystem with the option added. That depends on a private method.

Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
Lenguaje dominante
Python
Estrellas
1.2k
Forks
618
Merge medio
1 d 10 h
PR fusionados (30 d)
71

Preparar el entorno

  • Sin Dockerfile ni archivo de Docker Compose
  • Tiene una plantilla de pull request
  • Sin guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de apache/iceberg-python

Todos los issues de apache/iceberg-python

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.