Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Avro EnumReader.skip() does not advance the decoder

Aperta Adatta ai principianti
#4,006 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
1/5
Tempo stimato
Meno di un'ora
Idoneità per principianti
92/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
data

Direzione di ricerca

Inizia in pyiceberg/avro/resolver.py, in EnumReader.skip(), e segui il percorso di proiezione Avro per i campi enum omessi. Riproduci la lettura del manifest escludendo status dalla proiezione, quindi verifica che il successivo snapshot_id venga decodificato correttamente anziché come valore dell'enum. Il lavoro è completato quando il decoder avanza oltre il campo enum omesso.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

kind:bug
Apache Iceberg version

main (development)

Please describe the bug 🐞
Description

EnumReader.skip() in pyiceberg/avro/resolver.py currently does nothing:

def skip(self, decoder: BinaryDecoder) -> None:
    pass

When an enum field is omitted from the requested read schema, the Avro reader calls skip() for that field. Because the decoder is not advanced, the next field is read from the enum field's bytes.
This causes incorrect values when reading selected Avro records, including Iceberg manifest files.

Reproduction

Create an Iceberg table, append one row, and read its manifest with the status enum field projected out:

from tempfile import TemporaryDirectory

import pyarrow as pa

from pyiceberg.avro.file import AvroFile
from pyiceberg.catalog.memory import InMemoryCatalog
from pyiceberg.manifest import MANIFEST_ENTRY_SCHEMAS, ManifestEntryStatus
from pyiceberg.schema import Schema
from pyiceberg.types import IntegerType, NestedField

with TemporaryDirectory() as warehouse:
    # Use a temporary local warehouse so the example does not modify external data.
    catalog = InMemoryCatalog("bug-simulation", warehouse=warehouse)
    catalog.create_namespace("demo")

    # Create a simple Iceberg table with two required integer columns.
    table = catalog.create_table(
        "demo.events",
        schema=Schema(
            NestedField(1, "id", IntegerType(), required=True),
            NestedField(2, "value", IntegerType(), required=True),
        ),
    )

    # Build a PyArrow table whose types and nullability match the Iceberg schema.
    data = pa.Table.from_pylist(
        [{"id": 1, "value": 123}],
        schema=pa.schema(
            [
                pa.field("id", pa.int32(), nullable=False),
                pa.field("value", pa.int32(), nullable=False),
            ]
        ),
    )
    # Write the data file and commit a snapshot containing a manifest.
    table.append(data)

    # Find the manifest generated by the append operation.
    snapshot = table.current_snapshot()
    manifest = snapshot.manifests(table.io)[0]

    # The manifest schema starts with field ID 0, the status enum.
    file_schema = MANIFEST_ENTRY_SCHEMAS[2]

    # Build a projected schema that omits status but keeps the later fields.
    # This makes the Avro reader skip status before reading snapshot_id.
    projected_fields = []
    for field in file_schema.fields:
        if field.field_id != 0:
            projected_fields.append(field)

    projected_schema = Schema(*projected_fields)

    with AvroFile(
        table.io.new_input(manifest.manifest_path),
        read_schema=projected_schema,
        # Tell the resolver that field ID 0 should be converted to an enum.
        # The field is projected out, so EnumReader.skip() handles it.
        read_enums={0: ManifestEntryStatus},
    ) as reader:
        entries = list(reader)

    # Because status was projected out, the first returned field is snapshot_id.
    decoded_snapshot_id = entries[0][0]

    # the decoder is still positioned at status and returns 1.
    if decoded_snapshot_id != snapshot.snapshot_id:
        raise RuntimeError(
            f"Expected snapshot_id {snapshot.snapshot_id}, "
            f"got {decoded_snapshot_id}"
        )
Actual behavior

The decoded snapshot_id is incorrectly read as 1.
1 is the encoded manifest status value. This shows that the decoder did not skip the enum value before reading snapshot_id.

Expected behavior

The decoder should skip the enum value and decode the following snapshot_id correctly.

Proposed fix

Delegate skipping to the wrapped reader:

def skip(self, decoder: BinaryDecoder) -> None:
    self.reader.skip(decoder)
Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
Lingua principale
Python
Stelle
1.1k
Fork
589
Merge medio
2g 2h
PR unite (30g)
70

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/iceberg-python

Tutte le issue di apache/iceberg-python

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.