Metadata inspection APIs fail with struct.error after int→long / float→double type promotion
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 68/100
Hướng nghiên cứu
Tái hiện lỗi bằng script tự chứa, sau đó đọc pyiceberg/table/inspect.py tại InspectTable._get_files_from_manifest và pyiceberg/conversions.py tại from_bytes. So sánh độ dài bound đã mã hóa với các kiểu trường hiện tại cho int→long và float→double; được xem là hoàn tất khi files(), entries(), data_files() và all_files() trả về bound hoặc None mà không phát sinh struct.error.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Apache Iceberg version
0.11.0 (latest release)
Please describe the bug 🐞
Summary
After a type promotion that the spec allows (int → long, float → double), all metadata inspection APIs raise struct.error on any table that already contains data files written before the promotion:
| Promotion | inspect.files() |
inspect.entries() |
inspect.data_files() |
inspect.all_files() |
|---|---|---|---|---|
int → long |
❌ | ❌ | ❌ | ❌ |
float → double |
❌ | ❌ | ❌ | ❌ |
struct.error: unpack requires a buffer of 8 bytes
Since type promotion is not rewriting existing data files, the table stays in this state permanently — the whole metadata-inspection surface becomes unusable.
Reproduction
Self-contained, no cloud services or network required:
import tempfile, shutil, traceback
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, IntegerType, LongType, StringType
warehouse = tempfile.mkdtemp(prefix="iceberg_repro_")
catalog = SqlCatalog("repro", uri=f"sqlite:///{warehouse}/catalog.db",
warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")
tbl = catalog.create_table("ns.t", schema=Schema(
NestedField(1, "name", StringType(), required=False),
NestedField(2, "qty", IntegerType(), required=False),
))
# Write while the column is still `int` -> bounds are stored as 4-byte LE.
tbl.append(pa.Table.from_pylist(
[{"name": "a", "qty": 1}, {"name": "b", "qty": 2}],
schema=pa.schema([pa.field("name", pa.string(), nullable=True),
pa.field("qty", pa.int32(), nullable=True)]),
))
print("before promotion:", tbl.inspect.files().num_rows, "row(s) -> OK")
# Allowed promotion; existing data files and bounds are not rewritten.
with tbl.update_schema() as update:
update.update_column("qty", field_type=LongType())
tbl = catalog.load_table("ns.t")
try:
tbl.inspect.files()
except Exception:
traceback.print_exc()
shutil.rmtree(warehouse, ignore_errors=True)
Output
before promotion: 1 row(s) -> OK
Traceback (most recent call last):
...
File ".../pyiceberg/table/inspect.py", line 573, in _get_files_from_manifest
"lower_bound": from_bytes(field.field_type, lower_bound)
File ".../pyiceberg/conversions.py", line 337, in _
return _LONG_STRUCT.unpack(b)[0]
struct.error: unpack requires a buffer of 8 bytes
Replacing IntegerType()/LongType()/pa.int32() with FloatType()/DoubleType()/pa.float32() reproduces the same failure, as does calling entries(), data_files() or all_files() instead of files().
Expected behavior
inspect.files() and the other metadata tables should return the bounds (or omit/None them) rather than raising, on tables that have undergone a spec-allowed type promotion.
Analysis
InspectTable._get_files_from_manifest decodes lower_bounds / upper_bounds via from_bytes(field.field_type, ...), i.e. using the current schema type.
Per the spec, type promotion does not rewrite existing bounds, and the Avro manifest does not record which type was used to encode them. So after int → long, pre-existing files still carry 4-byte bounds while the current field type is long, and _LONG_STRUCT.unpack (8 bytes) fails.
As discussed on the dev list regarding type promotion in v3, implementations handle this by detecting the promotion from the encoded byte length rather than trusting the current schema type.
Workaround
Table.scan().plan_files() returns DataFile objects without decoding bounds, so file-level metadata (file_path, record_count, sort_order_id, spec_id, …) remains reachable:
tasks = list(tbl.scan().plan_files())
[(t.file.file_path, t.file.record_count) for t in tasks]
Environment
- pyiceberg 0.11.1 (latest release)
- pyarrow 25.0.0, SQLAlchemy 2.0.51
- Python 3.14.5, macOS (arm64)
Originally hit against an AWS S3 Tables table (Glue Iceberg REST catalog) whose int column had been promoted to bigint via Athena; the reproduction above shows it is catalog-independent.
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- Ngôn ngữ chính
- Python
- Star
- 1.1k
- Fork
- 589
- Merge trung bình
- 2 ngày 2 giờ
- Pull request đã merge (30 ngày)
- 70
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/iceberg-python
-
kind:bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
apache/iceberg-python#4006 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/iceberg-python#3996 ·
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validation Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
apache/iceberg-python#3979 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/iceberg-python#3885 ·
-
[Bug] PyArrowFileIO fails to propagate s3.ssl.ca-cert to pyarrow.fs.S3FileSystem tls_ca_file_path Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
apache/iceberg-python#3866 · 1 bình luận ·
Tất cả issue của apache/iceberg-python
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
anthropics/skills#1811 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
speaches-ai/speaches#678 ·
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
datalayer/mcp-compose#42 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
conda-forge/spacy-feedstock#177 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
UKGovernmentBEIS/inspect_evals#2523 ·