Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Metadata inspection APIs fail with struct.error after int→long / float→double type promotion

オープン
#3,744 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
68/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
静か
技術スタック
python
領域
databases

調査の方向性

自己完結型スクリプトで失敗を再現し、その後 pyiceberg/table/inspect.py の InspectTable._get_files_from_manifest と pyiceberg/conversions.py の from_bytes を読みます。int→long と float→double について、エンコードされた bound の長さを現在のフィールド型と比較します。files()、entries()、data_files()、all_files() が struct.error を発生させずに bounds または None を返せば完了です。

索引モデルが issue の本文から書いたものです。

説明

Apache Iceberg version

0.11.0 (latest release)

Please describe the bug 🐞
Summary

After a type promotion that the spec allows (intlong, floatdouble), all metadata inspection APIs raise struct.error on any table that already contains data files written before the promotion:

Promotion inspect.files() inspect.entries() inspect.data_files() inspect.all_files()
intlong
floatdouble
struct.error: unpack requires a buffer of 8 bytes

Since type promotion is not rewriting existing data files, the table stays in this state permanently — the whole metadata-inspection surface becomes unusable.

Reproduction

Self-contained, no cloud services or network required:

import tempfile, shutil, traceback
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, IntegerType, LongType, StringType

warehouse = tempfile.mkdtemp(prefix="iceberg_repro_")
catalog = SqlCatalog("repro", uri=f"sqlite:///{warehouse}/catalog.db",
                     warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")

tbl = catalog.create_table("ns.t", schema=Schema(
    NestedField(1, "name", StringType(), required=False),
    NestedField(2, "qty", IntegerType(), required=False),
))

# Write while the column is still `int` -> bounds are stored as 4-byte LE.
tbl.append(pa.Table.from_pylist(
    [{"name": "a", "qty": 1}, {"name": "b", "qty": 2}],
    schema=pa.schema([pa.field("name", pa.string(), nullable=True),
                      pa.field("qty", pa.int32(), nullable=True)]),
))
print("before promotion:", tbl.inspect.files().num_rows, "row(s) -> OK")

# Allowed promotion; existing data files and bounds are not rewritten.
with tbl.update_schema() as update:
    update.update_column("qty", field_type=LongType())

tbl = catalog.load_table("ns.t")
try:
    tbl.inspect.files()
except Exception:
    traceback.print_exc()

shutil.rmtree(warehouse, ignore_errors=True)

Output

before promotion: 1 row(s) -> OK
Traceback (most recent call last):
  ...
  File ".../pyiceberg/table/inspect.py", line 573, in _get_files_from_manifest
    "lower_bound": from_bytes(field.field_type, lower_bound)
  File ".../pyiceberg/conversions.py", line 337, in _
    return _LONG_STRUCT.unpack(b)[0]
struct.error: unpack requires a buffer of 8 bytes

Replacing IntegerType()/LongType()/pa.int32() with FloatType()/DoubleType()/pa.float32() reproduces the same failure, as does calling entries(), data_files() or all_files() instead of files().

Expected behavior

inspect.files() and the other metadata tables should return the bounds (or omit/None them) rather than raising, on tables that have undergone a spec-allowed type promotion.

Analysis

InspectTable._get_files_from_manifest decodes lower_bounds / upper_bounds via from_bytes(field.field_type, ...), i.e. using the current schema type.

Per the spec, type promotion does not rewrite existing bounds, and the Avro manifest does not record which type was used to encode them. So after intlong, pre-existing files still carry 4-byte bounds while the current field type is long, and _LONG_STRUCT.unpack (8 bytes) fails.

As discussed on the dev list regarding type promotion in v3, implementations handle this by detecting the promotion from the encoded byte length rather than trusting the current schema type.

Workaround

Table.scan().plan_files() returns DataFile objects without decoding bounds, so file-level metadata (file_path, record_count, sort_order_id, spec_id, …) remains reachable:

tasks = list(tbl.scan().plan_files())
[(t.file.file_path, t.file.record_count) for t in tasks]
Environment
  • pyiceberg 0.11.1 (latest release)
  • pyarrow 25.0.0, SQLAlchemy 2.0.51
  • Python 3.14.5, macOS (arm64)

Originally hit against an AWS S3 Tables table (Glue Iceberg REST catalog) whose int column had been promoted to bigint via Athena; the reproduction above shows it is catalog-independent.

Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
主要言語
Python
スター
1.1k
フォーク
589
平均マージ
2日 2時間
マージ済み PR(30日)
70

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

apache/iceberg-python のほかの issue

apache/iceberg-python の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。