Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

PyArrowFileIO: every small S3 write is a 3-request multipart upload; expose allow_delayed_open

Open Beginner friendly
#4,093 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
70/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Active
Tech stack
aws, python
Domain
cloud, data

Research direction

The issue names _initialize_s3_fs in PyArrowFileIO as the place where the S3 filesystem is built from a fixed set of properties; start there and see how existing s3.* properties are read and passed on. Add an s3.allow-delayed-open property, pass allow_delayed_open only when the installed pyarrow is 21 or later, and default it to true as the issue proposes. Done when small writes go out as one PutObject and the option is absent on older pyarrow.

Written by the indexing model from the issue text.

Description

Apache Iceberg version

0.12.0 (also on main)

Please describe the bug 🐞

PyArrowFileIO writes every S3 object as a multipart upload, however small it is. pyarrow's S3FileSystem.open_output_stream starts a multipart upload as soon as the stream opens (apache/arrow#51029). Each metadata.json, manifest list, manifest, version-hint.text and small data file therefore costs three requests: CreateMultipartUpload, UploadPart and CompleteMultipartUpload. All three are billed as writes, even for a 1-byte object.

pyarrow's S3FileSystem has an option for this, allow_delayed_open, since pyarrow 21. With it set, a stream that closes before reaching a part's size is sent as a single PutObject, and a larger one is still a multipart upload. _initialize_s3_fs builds the filesystem from a fixed set of properties, though, so there's no way to set the option through FileIO properties.

The cost is real for a table that commits often. Each commit writes several small metadata objects, so it makes about three times the write requests it needs.

Repro

Run against any S3-compatible endpoint (this was run against a local rustfs):

import pyarrow.fs as fs
from pyiceberg.io.pyarrow import PyArrowFileIO

fs.initialize_s3(fs.S3LogLevel.Debug)  # logs each request

io = PyArrowFileIO({
    "s3.endpoint": "http://127.0.0.1:9000",
    "s3.access-key-id": "...",
    "s3.secret-access-key": "...",
    "s3.region": "us-east-1",
})
with io.new_output("s3://bucket/version-hint.text").create(overwrite=True) as f:
    f.write(b"1")

The debug log shows three requests for the one byte:

POST /bucket/version-hint.text?uploads
PUT  /bucket/version-hint.text?partNumber=1&uploadId=...
POST /bucket/version-hint.text?uploadId=...

When the same write goes through an S3FileSystem built with allow_delayed_open=True, the log shows a single PUT /bucket/version-hint.text.

Proposal

Pass allow_delayed_open in _initialize_s3_fs, controlled by a FileIO property such as s3.allow-delayed-open. I'd suggest defaulting it to true, since it only changes how small objects are uploaded. pyiceberg supports pyarrow 18 and up, so the option would be passed only on pyarrow 21 or later.

The workaround today is to subclass PyArrowFileIO, override _initialize_s3_fs, and rebuild the filesystem with the option added. That depends on a private method.

Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
Dominant language
Python
Stars
1.2k
Forks
618
Avg merge
1d 18h
Merged PRs (30d)
70

Getting set up

  • No Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/iceberg-python

All issues in apache/iceberg-python

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.