IO: Consolidate PyArrow logic into io/pyarrow.py before decomposition

Open
#3,812 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
55/100
Issue type
Refactor
Clarity
Mostly clear
Activity status
Active
Tech stack
python
Domain
backend

Research direction

Start by comparing table/upsert_util.py with pyiceberg/io/pyarrow.py and reviewing the related discussion in #3737. Focus on PR A first; done means the upsert utility's PyArrow logic routes through io/pyarrow.py, with no behavior changes and all existing tests passing.

Written by the indexing model from the issue text.

Description

Summary

Before decomposing pyiceberg/io/pyarrow.py into focused submodules (#3737, #3738), we should consolidate PyArrow-specific logic that currently lives outside the module. This ensures all PyArrow calls route through a single boundary, making the subsequent split clean and enabling future engine substitution.

Motivation

Per discussion in #3737, @rambleraptor noted that the first useful step is ensuring no PyArrow logic occurs outside pyarrow.py. Currently several modules import pyarrow directly and implement compute logic inline rather than delegating through pyiceberg.io.pyarrow.

When we later introduce a ComputeEngine protocol, any PyArrow logic outside the module boundary bypasses the protocol and prevents clean substitution.

Audit

Grepped pyiceberg/ (excluding io/pyarrow.py and tests) for runtime import pyarrow statements (both top-level and inline). Excluded TYPE_CHECKING-only imports since those have no runtime dependency.

Location What it does Action
table/upsert_util.py PyArrow table joins, group_by, compute, cast, take Absorb
table/inspect.py Builds pa.schema + pa.Table.from_pylist for metadata inspection TBD
transforms.py pyarrow_transform() dispatch on pa.Array/ChunkedArray TBD
table/__init__.py Entry points accept pa.Table, delegate to io.pyarrow Leave
table/deletion_vector.py Single pa.chunked_array() call Leave
catalog/__init__.py Delegates to io.pyarrow for schema conversion Leave

Plan

One PR per absorption. Each is a pure refactor: move code into io/pyarrow.py, have the caller import from pyiceberg.io.pyarrow instead of pyarrow directly. No behavior change, all existing tests pass unchanged.

  • PR A: Absorb table/upsert_util.py PyArrow logic
  • PR B: table/inspect.py (pending discussion)
  • PR C: transforms.py (pending discussion)

Related

  • #3737 - Decompose io/pyarrow.py into focused modules
  • #3738 - Extract PyArrowFileIO (first decomposition step)
  • #3715 / #3716 - Previous pluggable backend attempt (rejected as too large)
Dominant language
Python
Stars
1.1k
Forks
589
Avg merge
2d 4h
Merged PRs (30d)
72

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/iceberg-python

All issues in apache/iceberg-python

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.