Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

Avro EnumReader.skip() does not advance the decoder

Fermée Adaptée aux débutants
#4,006 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Les mainteneurs répondent en général sous 1 jour

Personne n'a encore pris cette issue.

Évaluation

Difficulté
1/5
Temps estimé
Moins d'une heure
Accessibilité débutants
92/100
Type d'issue
Bug
Clarté
Clairement spécifiée
Activité
Active
Stack technique
python
Domaine
data

Piste de recherche

Commencez dans pyiceberg/avro/resolver.py, au niveau de EnumReader.skip(), et suivez le chemin de projection Avro pour les champs enum omis. Reproduisez la lecture du manifest en excluant status de la projection, puis vérifiez que le snapshot_id suivant est décodé correctement plutôt que comme la valeur de l'enum. Le travail est terminé lorsque le décodeur avance au-delà du champ enum omis.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

kind:bug
Apache Iceberg version

main (development)

Please describe the bug 🐞
Description

EnumReader.skip() in pyiceberg/avro/resolver.py currently does nothing:

def skip(self, decoder: BinaryDecoder) -> None:
    pass

When an enum field is omitted from the requested read schema, the Avro reader calls skip() for that field. Because the decoder is not advanced, the next field is read from the enum field's bytes.
This causes incorrect values when reading selected Avro records, including Iceberg manifest files.

Reproduction

Create an Iceberg table, append one row, and read its manifest with the status enum field projected out:

from tempfile import TemporaryDirectory

import pyarrow as pa

from pyiceberg.avro.file import AvroFile
from pyiceberg.catalog.memory import InMemoryCatalog
from pyiceberg.manifest import MANIFEST_ENTRY_SCHEMAS, ManifestEntryStatus
from pyiceberg.schema import Schema
from pyiceberg.types import IntegerType, NestedField

with TemporaryDirectory() as warehouse:
    # Use a temporary local warehouse so the example does not modify external data.
    catalog = InMemoryCatalog("bug-simulation", warehouse=warehouse)
    catalog.create_namespace("demo")

    # Create a simple Iceberg table with two required integer columns.
    table = catalog.create_table(
        "demo.events",
        schema=Schema(
            NestedField(1, "id", IntegerType(), required=True),
            NestedField(2, "value", IntegerType(), required=True),
        ),
    )

    # Build a PyArrow table whose types and nullability match the Iceberg schema.
    data = pa.Table.from_pylist(
        [{"id": 1, "value": 123}],
        schema=pa.schema(
            [
                pa.field("id", pa.int32(), nullable=False),
                pa.field("value", pa.int32(), nullable=False),
            ]
        ),
    )
    # Write the data file and commit a snapshot containing a manifest.
    table.append(data)

    # Find the manifest generated by the append operation.
    snapshot = table.current_snapshot()
    manifest = snapshot.manifests(table.io)[0]

    # The manifest schema starts with field ID 0, the status enum.
    file_schema = MANIFEST_ENTRY_SCHEMAS[2]

    # Build a projected schema that omits status but keeps the later fields.
    # This makes the Avro reader skip status before reading snapshot_id.
    projected_fields = []
    for field in file_schema.fields:
        if field.field_id != 0:
            projected_fields.append(field)

    projected_schema = Schema(*projected_fields)

    with AvroFile(
        table.io.new_input(manifest.manifest_path),
        read_schema=projected_schema,
        # Tell the resolver that field ID 0 should be converted to an enum.
        # The field is projected out, so EnumReader.skip() handles it.
        read_enums={0: ManifestEntryStatus},
    ) as reader:
        entries = list(reader)

    # Because status was projected out, the first returned field is snapshot_id.
    decoded_snapshot_id = entries[0][0]

    # the decoder is still positioned at status and returns 1.
    if decoded_snapshot_id != snapshot.snapshot_id:
        raise RuntimeError(
            f"Expected snapshot_id {snapshot.snapshot_id}, "
            f"got {decoded_snapshot_id}"
        )
Actual behavior

The decoded snapshot_id is incorrectly read as 1.
1 is the encoded manifest status value. This shows that the decoder did not skip the enum value before reading snapshot_id.

Expected behavior

The decoder should skip the enum value and decode the following snapshot_id correctly.

Proposed fix

Delegate skipping to the wrapped reader:

def skip(self, decoder: BinaryDecoder) -> None:
    self.reader.skip(decoder)
Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
Langage dominant
Python
Étoiles
1.2k
Forks
618
Merge moyen
1 j 18 h
PR mergées (30 j)
70

Préparer son environnement

  • Aucun Dockerfile ni fichier Docker Compose
  • Propose un modèle de pull request
  • Aucun guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de apache/iceberg-python

Toutes les issues de apache/iceberg-python

Issues similaires

Plus d'issues Python

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.