Avro EnumReader.skip() does not advance the decoder
还没有人认领这个 Issue。
评估
调研方向
从 pyiceberg/avro/resolver.py 中的 EnumReader.skip() 开始,跟踪被省略 enum 字段的 Avro 投影路径。重现将 status 排除在投影之外的 manifest 读取,然后验证后续的 snapshot_id 是否被正确解码,而不是被解码为 enum 值。当 decoder 越过被省略的 enum 字段继续前进时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
Apache Iceberg version
main (development)
Please describe the bug 🐞
Description
EnumReader.skip() in pyiceberg/avro/resolver.py currently does nothing:
def skip(self, decoder: BinaryDecoder) -> None:
pass
When an enum field is omitted from the requested read schema, the Avro reader calls skip() for that field. Because the decoder is not advanced, the next field is read from the enum field's bytes.
This causes incorrect values when reading selected Avro records, including Iceberg manifest files.
Reproduction
Create an Iceberg table, append one row, and read its manifest with the status enum field projected out:
from tempfile import TemporaryDirectory
import pyarrow as pa
from pyiceberg.avro.file import AvroFile
from pyiceberg.catalog.memory import InMemoryCatalog
from pyiceberg.manifest import MANIFEST_ENTRY_SCHEMAS, ManifestEntryStatus
from pyiceberg.schema import Schema
from pyiceberg.types import IntegerType, NestedField
with TemporaryDirectory() as warehouse:
# Use a temporary local warehouse so the example does not modify external data.
catalog = InMemoryCatalog("bug-simulation", warehouse=warehouse)
catalog.create_namespace("demo")
# Create a simple Iceberg table with two required integer columns.
table = catalog.create_table(
"demo.events",
schema=Schema(
NestedField(1, "id", IntegerType(), required=True),
NestedField(2, "value", IntegerType(), required=True),
),
)
# Build a PyArrow table whose types and nullability match the Iceberg schema.
data = pa.Table.from_pylist(
[{"id": 1, "value": 123}],
schema=pa.schema(
[
pa.field("id", pa.int32(), nullable=False),
pa.field("value", pa.int32(), nullable=False),
]
),
)
# Write the data file and commit a snapshot containing a manifest.
table.append(data)
# Find the manifest generated by the append operation.
snapshot = table.current_snapshot()
manifest = snapshot.manifests(table.io)[0]
# The manifest schema starts with field ID 0, the status enum.
file_schema = MANIFEST_ENTRY_SCHEMAS[2]
# Build a projected schema that omits status but keeps the later fields.
# This makes the Avro reader skip status before reading snapshot_id.
projected_fields = []
for field in file_schema.fields:
if field.field_id != 0:
projected_fields.append(field)
projected_schema = Schema(*projected_fields)
with AvroFile(
table.io.new_input(manifest.manifest_path),
read_schema=projected_schema,
# Tell the resolver that field ID 0 should be converted to an enum.
# The field is projected out, so EnumReader.skip() handles it.
read_enums={0: ManifestEntryStatus},
) as reader:
entries = list(reader)
# Because status was projected out, the first returned field is snapshot_id.
decoded_snapshot_id = entries[0][0]
# the decoder is still positioned at status and returns 1.
if decoded_snapshot_id != snapshot.snapshot_id:
raise RuntimeError(
f"Expected snapshot_id {snapshot.snapshot_id}, "
f"got {decoded_snapshot_id}"
)
Actual behavior
The decoded snapshot_id is incorrectly read as 1.
1 is the encoded manifest status value. This shows that the decoder did not skip the enum value before reading snapshot_id.
Expected behavior
The decoder should skip the enum value and decode the following snapshot_id correctly.
Proposed fix
Delegate skipping to the wrapped reader:
def skip(self, decoder: BinaryDecoder) -> None:
self.reader.skip(decoder)
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- 主要语言
- Python
- 星标
- 1.1k
- 派生
- 589
- 平均合并
- 2 天 2 小时
- 30 天内合并 PR
- 70
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/iceberg-python 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 78/100
apache/iceberg-python#3996 ·
-
bug
难度 2/5 1-3 小时 新手友好度 72/100
apache/iceberg-python#3979 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
apache/iceberg-python#3885 ·
-
[Bug] PyArrowFileIO fails to propagate s3.ssl.ca-cert to pyarrow.fs.S3FileSystem tls_ca_file_path 未关闭
难度 2/5 1-3 小时 新手友好度 76/100
apache/iceberg-python#3866 · 1 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
apache/iceberg-python#3836 · 1 条评论 ·
查看 apache/iceberg-python 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 75/100
anthropics/skills#1811 · 1 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
speaches-ai/speaches#678 ·
-
bug
难度 2/5 1-3 小时 新手友好度 75/100
datalayer/mcp-compose#42 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
conda-forge/spacy-feedstock#177 ·
-
难度 2/5 1-3 小时 新手友好度 70/100
UKGovernmentBEIS/inspect_evals#2523 ·