Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

PyArrowFileIO: every small S3 write is a 3-request multipart upload; expose allow_delayed_open

Aperta Adatta ai principianti
#4,093 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
70/100
Tipo di issue
Funzionalità
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
aws, python
Ambito
cloud, data

Direzione di ricerca

Inizia da _initialize_s3_fs in PyArrowFileIO, dove il filesystem S3 viene costruito da un insieme fisso di proprietà. Verifica come vengono letti e passati le proprietà s3.* esistenti. Aggiungi una proprietà s3.allow-delayed-open, passa allow_delayed_open solo se il pyarrow installato è la 21 o successiva, e imposta il valore predefinito a true come proposto nella issue. È completato quando le scritture piccole vengono inviate come un unico PutObject e l'opzione è assente con pyarrow più vecchi.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Apache Iceberg version

0.12.0 (also on main)

Please describe the bug 🐞

PyArrowFileIO writes every S3 object as a multipart upload, however small it is. pyarrow's S3FileSystem.open_output_stream starts a multipart upload as soon as the stream opens (apache/arrow#51029). Each metadata.json, manifest list, manifest, version-hint.text and small data file therefore costs three requests: CreateMultipartUpload, UploadPart and CompleteMultipartUpload. All three are billed as writes, even for a 1-byte object.

pyarrow's S3FileSystem has an option for this, allow_delayed_open, since pyarrow 21. With it set, a stream that closes before reaching a part's size is sent as a single PutObject, and a larger one is still a multipart upload. _initialize_s3_fs builds the filesystem from a fixed set of properties, though, so there's no way to set the option through FileIO properties.

The cost is real for a table that commits often. Each commit writes several small metadata objects, so it makes about three times the write requests it needs.

Repro

Run against any S3-compatible endpoint (this was run against a local rustfs):

import pyarrow.fs as fs
from pyiceberg.io.pyarrow import PyArrowFileIO

fs.initialize_s3(fs.S3LogLevel.Debug)  # logs each request

io = PyArrowFileIO({
    "s3.endpoint": "http://127.0.0.1:9000",
    "s3.access-key-id": "...",
    "s3.secret-access-key": "...",
    "s3.region": "us-east-1",
})
with io.new_output("s3://bucket/version-hint.text").create(overwrite=True) as f:
    f.write(b"1")

The debug log shows three requests for the one byte:

POST /bucket/version-hint.text?uploads
PUT  /bucket/version-hint.text?partNumber=1&uploadId=...
POST /bucket/version-hint.text?uploadId=...

When the same write goes through an S3FileSystem built with allow_delayed_open=True, the log shows a single PUT /bucket/version-hint.text.

Proposal

Pass allow_delayed_open in _initialize_s3_fs, controlled by a FileIO property such as s3.allow-delayed-open. I'd suggest defaulting it to true, since it only changes how small objects are uploaded. pyiceberg supports pyarrow 18 and up, so the option would be passed only on pyarrow 21 or later.

The workaround today is to subclass PyArrowFileIO, override _initialize_s3_fs, and rebuild the filesystem with the option added. That depends on a private method.

Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
Lingua principale
Python
Stelle
1.2k
Fork
618
Merge medio
1g 10h
PR unite (30g)
71

Preparare l'ambiente

  • Nessun Dockerfile né file Docker Compose
  • Ha un modello di pull request
  • Nessuna guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/iceberg-python

Tutte le issue di apache/iceberg-python

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.