Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

PyArrowFileIO: every small S3 write is a 3-request multipart upload; expose allow_delayed_open

Offen Anfängerfreundlich
#4,093 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Maintainer antworten meist innerhalb von 1 Tag

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
2/5
Geschätzter Aufwand
1-3 Stunden
Anfängerfreundlichkeit
70/100
Issue-Typ
Feature
Klarheit
Klar beschrieben
Aktivitätsstatus
Aktiv
Tech-Stack
aws, python
Bereich
cloud, data

Rechercherichtung

Beginne bei _initialize_s3_fs in PyArrowFileIO, wo das S3-Dateisystem aus einer festen Menge von Eigenschaften erstellt wird. Prüfe, wie bestehende s3.*-Eigenschaften gelesen und weitergereicht werden. Füge eine Eigenschaft s3.allow-delayed-open hinzu, übergib allow_delayed_open nur dann, wenn das installierte pyarrow Version 21 oder neuer ist, und setze den Standardwert auf true, wie im Issue vorgeschlagen. Fertig ist es, wenn kleine Schreibvorgänge als ein einziges PutObject gesendet werden und die Option bei älterem pyarrow fehlt.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Apache Iceberg version

0.12.0 (also on main)

Please describe the bug 🐞

PyArrowFileIO writes every S3 object as a multipart upload, however small it is. pyarrow's S3FileSystem.open_output_stream starts a multipart upload as soon as the stream opens (apache/arrow#51029). Each metadata.json, manifest list, manifest, version-hint.text and small data file therefore costs three requests: CreateMultipartUpload, UploadPart and CompleteMultipartUpload. All three are billed as writes, even for a 1-byte object.

pyarrow's S3FileSystem has an option for this, allow_delayed_open, since pyarrow 21. With it set, a stream that closes before reaching a part's size is sent as a single PutObject, and a larger one is still a multipart upload. _initialize_s3_fs builds the filesystem from a fixed set of properties, though, so there's no way to set the option through FileIO properties.

The cost is real for a table that commits often. Each commit writes several small metadata objects, so it makes about three times the write requests it needs.

Repro

Run against any S3-compatible endpoint (this was run against a local rustfs):

import pyarrow.fs as fs
from pyiceberg.io.pyarrow import PyArrowFileIO

fs.initialize_s3(fs.S3LogLevel.Debug)  # logs each request

io = PyArrowFileIO({
    "s3.endpoint": "http://127.0.0.1:9000",
    "s3.access-key-id": "...",
    "s3.secret-access-key": "...",
    "s3.region": "us-east-1",
})
with io.new_output("s3://bucket/version-hint.text").create(overwrite=True) as f:
    f.write(b"1")

The debug log shows three requests for the one byte:

POST /bucket/version-hint.text?uploads
PUT  /bucket/version-hint.text?partNumber=1&uploadId=...
POST /bucket/version-hint.text?uploadId=...

When the same write goes through an S3FileSystem built with allow_delayed_open=True, the log shows a single PUT /bucket/version-hint.text.

Proposal

Pass allow_delayed_open in _initialize_s3_fs, controlled by a FileIO property such as s3.allow-delayed-open. I'd suggest defaulting it to true, since it only changes how small objects are uploaded. pyiceberg supports pyarrow 18 and up, so the option would be passed only on pyarrow 21 or later.

The workaround today is to subclass PyArrowFileIO, override _initialize_s3_fs, and rebuild the filesystem with the option added. That depends on a private method.

Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
Vorherrschende Sprache
Python
Sterne
1.2k
Forks
618
Ø Merge
1 T. 10 Std.
Gemergte PRs (30 T.)
71

Entwicklungsumgebung

  • Kein Dockerfile und keine Docker-Compose-Datei
  • Hat eine Pull-Request-Vorlage
  • Kein Beitragsleitfaden

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus apache/iceberg-python

Alle Issues in apache/iceberg-python

Ähnliche Issues

Weitere Issues zu Python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.