Metadata inspection APIs fail with struct.error after int→long / float→double type promotion
还没有人认领这个 Issue。
评估
调研方向
使用自包含脚本复现该失败,然后阅读 pyiceberg/table/inspect.py 中的 InspectTable._get_files_from_manifest 和 pyiceberg/conversions.py 中的 from_bytes。比较 int→long 和 float→double 的编码 bound 长度与当前字段类型;当 files()、entries()、data_files() 和 all_files() 返回 bound 或 None 且不引发 struct.error 时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
Apache Iceberg version
0.11.0 (latest release)
Please describe the bug 🐞
Summary
After a type promotion that the spec allows (int → long, float → double), all metadata inspection APIs raise struct.error on any table that already contains data files written before the promotion:
| Promotion | inspect.files() |
inspect.entries() |
inspect.data_files() |
inspect.all_files() |
|---|---|---|---|---|
int → long |
❌ | ❌ | ❌ | ❌ |
float → double |
❌ | ❌ | ❌ | ❌ |
struct.error: unpack requires a buffer of 8 bytes
Since type promotion is not rewriting existing data files, the table stays in this state permanently — the whole metadata-inspection surface becomes unusable.
Reproduction
Self-contained, no cloud services or network required:
import tempfile, shutil, traceback
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, IntegerType, LongType, StringType
warehouse = tempfile.mkdtemp(prefix="iceberg_repro_")
catalog = SqlCatalog("repro", uri=f"sqlite:///{warehouse}/catalog.db",
warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")
tbl = catalog.create_table("ns.t", schema=Schema(
NestedField(1, "name", StringType(), required=False),
NestedField(2, "qty", IntegerType(), required=False),
))
# Write while the column is still `int` -> bounds are stored as 4-byte LE.
tbl.append(pa.Table.from_pylist(
[{"name": "a", "qty": 1}, {"name": "b", "qty": 2}],
schema=pa.schema([pa.field("name", pa.string(), nullable=True),
pa.field("qty", pa.int32(), nullable=True)]),
))
print("before promotion:", tbl.inspect.files().num_rows, "row(s) -> OK")
# Allowed promotion; existing data files and bounds are not rewritten.
with tbl.update_schema() as update:
update.update_column("qty", field_type=LongType())
tbl = catalog.load_table("ns.t")
try:
tbl.inspect.files()
except Exception:
traceback.print_exc()
shutil.rmtree(warehouse, ignore_errors=True)
Output
before promotion: 1 row(s) -> OK
Traceback (most recent call last):
...
File ".../pyiceberg/table/inspect.py", line 573, in _get_files_from_manifest
"lower_bound": from_bytes(field.field_type, lower_bound)
File ".../pyiceberg/conversions.py", line 337, in _
return _LONG_STRUCT.unpack(b)[0]
struct.error: unpack requires a buffer of 8 bytes
Replacing IntegerType()/LongType()/pa.int32() with FloatType()/DoubleType()/pa.float32() reproduces the same failure, as does calling entries(), data_files() or all_files() instead of files().
Expected behavior
inspect.files() and the other metadata tables should return the bounds (or omit/None them) rather than raising, on tables that have undergone a spec-allowed type promotion.
Analysis
InspectTable._get_files_from_manifest decodes lower_bounds / upper_bounds via from_bytes(field.field_type, ...), i.e. using the current schema type.
Per the spec, type promotion does not rewrite existing bounds, and the Avro manifest does not record which type was used to encode them. So after int → long, pre-existing files still carry 4-byte bounds while the current field type is long, and _LONG_STRUCT.unpack (8 bytes) fails.
As discussed on the dev list regarding type promotion in v3, implementations handle this by detecting the promotion from the encoded byte length rather than trusting the current schema type.
Workaround
Table.scan().plan_files() returns DataFile objects without decoding bounds, so file-level metadata (file_path, record_count, sort_order_id, spec_id, …) remains reachable:
tasks = list(tbl.scan().plan_files())
[(t.file.file_path, t.file.record_count) for t in tasks]
Environment
- pyiceberg 0.11.1 (latest release)
- pyarrow 25.0.0, SQLAlchemy 2.0.51
- Python 3.14.5, macOS (arm64)
Originally hit against an AWS S3 Tables table (Glue Iceberg REST catalog) whose int column had been promoted to bigint via Athena; the reproduction above shows it is catalog-independent.
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- 主要语言
- Python
- 星标
- 1.1k
- 派生
- 589
- 平均合并
- 1 天 20 小时
- 30 天内合并 PR
- 68
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/iceberg-python 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 70/100
apache/iceberg-python#4010 · 1 个 reaction ·
-
kind:bug
难度 1/5 1 小时以内 新手友好度 92/100
apache/iceberg-python#4006 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
apache/iceberg-python#3996 ·
-
bug
难度 2/5 1-3 小时 新手友好度 72/100
apache/iceberg-python#3979 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
apache/iceberg-python#3885 ·
查看 apache/iceberg-python 的全部 Issue
相似的 Issue
-
bug
难度 2/5 1-3 小时 新手友好度 75/100
xinnan-tech/xiaozhi-fde-talk#263 ·
-
rules
难度 1/5 1 小时以内 新手友好度 90/100
-
难度 2/5 1-3 小时 新手友好度 70/100
huggingface/Repo2RLEnv#163 · 1 条评论 ·
-
难度 1/5 1 小时以内 新手友好度 95/100
huggingface/sentence-transformers#4074 ·
-
comp/dashboard invalid P3
难度 2/5 1-3 小时 新手友好度 70/100
NousResearch/hermes-agent#121143 ·