Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Avro EnumReader.skip() does not advance the decoder

オープン 初心者向け
#4,006 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
1/5
見積もり時間
1時間未満
初心者へのやさしさ
92/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python
領域
data

調査の方向性

pyiceberg/avro/resolver.py の EnumReader.skip() から始め、欠落した enum フィールドに対する Avro の projection パスを追跡します。status を projection から除外した状態で manifest の読み取りを再現し、その後続の snapshot_id が enum 値としてではなく正しくデコードされることを確認します。decoder が欠落した enum フィールドを越えて進むことを確認できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

kind:bug
Apache Iceberg version

main (development)

Please describe the bug 🐞
Description

EnumReader.skip() in pyiceberg/avro/resolver.py currently does nothing:

def skip(self, decoder: BinaryDecoder) -> None:
    pass

When an enum field is omitted from the requested read schema, the Avro reader calls skip() for that field. Because the decoder is not advanced, the next field is read from the enum field's bytes.
This causes incorrect values when reading selected Avro records, including Iceberg manifest files.

Reproduction

Create an Iceberg table, append one row, and read its manifest with the status enum field projected out:

from tempfile import TemporaryDirectory

import pyarrow as pa

from pyiceberg.avro.file import AvroFile
from pyiceberg.catalog.memory import InMemoryCatalog
from pyiceberg.manifest import MANIFEST_ENTRY_SCHEMAS, ManifestEntryStatus
from pyiceberg.schema import Schema
from pyiceberg.types import IntegerType, NestedField

with TemporaryDirectory() as warehouse:
    # Use a temporary local warehouse so the example does not modify external data.
    catalog = InMemoryCatalog("bug-simulation", warehouse=warehouse)
    catalog.create_namespace("demo")

    # Create a simple Iceberg table with two required integer columns.
    table = catalog.create_table(
        "demo.events",
        schema=Schema(
            NestedField(1, "id", IntegerType(), required=True),
            NestedField(2, "value", IntegerType(), required=True),
        ),
    )

    # Build a PyArrow table whose types and nullability match the Iceberg schema.
    data = pa.Table.from_pylist(
        [{"id": 1, "value": 123}],
        schema=pa.schema(
            [
                pa.field("id", pa.int32(), nullable=False),
                pa.field("value", pa.int32(), nullable=False),
            ]
        ),
    )
    # Write the data file and commit a snapshot containing a manifest.
    table.append(data)

    # Find the manifest generated by the append operation.
    snapshot = table.current_snapshot()
    manifest = snapshot.manifests(table.io)[0]

    # The manifest schema starts with field ID 0, the status enum.
    file_schema = MANIFEST_ENTRY_SCHEMAS[2]

    # Build a projected schema that omits status but keeps the later fields.
    # This makes the Avro reader skip status before reading snapshot_id.
    projected_fields = []
    for field in file_schema.fields:
        if field.field_id != 0:
            projected_fields.append(field)

    projected_schema = Schema(*projected_fields)

    with AvroFile(
        table.io.new_input(manifest.manifest_path),
        read_schema=projected_schema,
        # Tell the resolver that field ID 0 should be converted to an enum.
        # The field is projected out, so EnumReader.skip() handles it.
        read_enums={0: ManifestEntryStatus},
    ) as reader:
        entries = list(reader)

    # Because status was projected out, the first returned field is snapshot_id.
    decoded_snapshot_id = entries[0][0]

    # the decoder is still positioned at status and returns 1.
    if decoded_snapshot_id != snapshot.snapshot_id:
        raise RuntimeError(
            f"Expected snapshot_id {snapshot.snapshot_id}, "
            f"got {decoded_snapshot_id}"
        )
Actual behavior

The decoded snapshot_id is incorrectly read as 1.
1 is the encoded manifest status value. This shows that the decoder did not skip the enum value before reading snapshot_id.

Expected behavior

The decoder should skip the enum value and decode the following snapshot_id correctly.

Proposed fix

Delegate skipping to the wrapped reader:

def skip(self, decoder: BinaryDecoder) -> None:
    self.reader.skip(decoder)
Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
主要言語
Python
スター
1.1k
フォーク
589
平均マージ
2日 2時間
マージ済み PR(30日)
70

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

apache/iceberg-python のほかの issue

apache/iceberg-python の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。