Avro EnumReader.skip() does not advance the decoder
Les mainteneurs répondent en général sous 1 jour
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 1/5
- Temps estimé
- Moins d'une heure
- Accessibilité débutants
- 92/100
Piste de recherche
Commencez dans pyiceberg/avro/resolver.py, au niveau de EnumReader.skip(), et suivez le chemin de projection Avro pour les champs enum omis. Reproduisez la lecture du manifest en excluant status de la projection, puis vérifiez que le snapshot_id suivant est décodé correctement plutôt que comme la valeur de l'enum. Le travail est terminé lorsque le décodeur avance au-delà du champ enum omis.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Apache Iceberg version
main (development)
Please describe the bug 🐞
Description
EnumReader.skip() in pyiceberg/avro/resolver.py currently does nothing:
def skip(self, decoder: BinaryDecoder) -> None:
pass
When an enum field is omitted from the requested read schema, the Avro reader calls skip() for that field. Because the decoder is not advanced, the next field is read from the enum field's bytes.
This causes incorrect values when reading selected Avro records, including Iceberg manifest files.
Reproduction
Create an Iceberg table, append one row, and read its manifest with the status enum field projected out:
from tempfile import TemporaryDirectory
import pyarrow as pa
from pyiceberg.avro.file import AvroFile
from pyiceberg.catalog.memory import InMemoryCatalog
from pyiceberg.manifest import MANIFEST_ENTRY_SCHEMAS, ManifestEntryStatus
from pyiceberg.schema import Schema
from pyiceberg.types import IntegerType, NestedField
with TemporaryDirectory() as warehouse:
# Use a temporary local warehouse so the example does not modify external data.
catalog = InMemoryCatalog("bug-simulation", warehouse=warehouse)
catalog.create_namespace("demo")
# Create a simple Iceberg table with two required integer columns.
table = catalog.create_table(
"demo.events",
schema=Schema(
NestedField(1, "id", IntegerType(), required=True),
NestedField(2, "value", IntegerType(), required=True),
),
)
# Build a PyArrow table whose types and nullability match the Iceberg schema.
data = pa.Table.from_pylist(
[{"id": 1, "value": 123}],
schema=pa.schema(
[
pa.field("id", pa.int32(), nullable=False),
pa.field("value", pa.int32(), nullable=False),
]
),
)
# Write the data file and commit a snapshot containing a manifest.
table.append(data)
# Find the manifest generated by the append operation.
snapshot = table.current_snapshot()
manifest = snapshot.manifests(table.io)[0]
# The manifest schema starts with field ID 0, the status enum.
file_schema = MANIFEST_ENTRY_SCHEMAS[2]
# Build a projected schema that omits status but keeps the later fields.
# This makes the Avro reader skip status before reading snapshot_id.
projected_fields = []
for field in file_schema.fields:
if field.field_id != 0:
projected_fields.append(field)
projected_schema = Schema(*projected_fields)
with AvroFile(
table.io.new_input(manifest.manifest_path),
read_schema=projected_schema,
# Tell the resolver that field ID 0 should be converted to an enum.
# The field is projected out, so EnumReader.skip() handles it.
read_enums={0: ManifestEntryStatus},
) as reader:
entries = list(reader)
# Because status was projected out, the first returned field is snapshot_id.
decoded_snapshot_id = entries[0][0]
# the decoder is still positioned at status and returns 1.
if decoded_snapshot_id != snapshot.snapshot_id:
raise RuntimeError(
f"Expected snapshot_id {snapshot.snapshot_id}, "
f"got {decoded_snapshot_id}"
)
Actual behavior
The decoded snapshot_id is incorrectly read as 1.
1 is the encoded manifest status value. This shows that the decoder did not skip the enum value before reading snapshot_id.
Expected behavior
The decoder should skip the enum value and decode the following snapshot_id correctly.
Proposed fix
Delegate skipping to the wrapped reader:
def skip(self, decoder: BinaryDecoder) -> None:
self.reader.skip(decoder)
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- Langage dominant
- Python
- Étoiles
- 1.2k
- Forks
- 618
- Merge moyen
- 1 j 18 h
- PR mergées (30 j)
- 70
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Propose un modèle de pull request
- Aucun guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de apache/iceberg-python
-
PyArrowFileIO: every small S3 write is a 3-request multipart upload; expose allow_delayed_openOuverte
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
apache/iceberg-python#4093 ·
Les mainteneurs répondent en général sous 1 jour
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's responsePeut-être pris @Soumo-git-hub l’a pris il y a 3 jours. Ouvertekind:bug
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100
apache/iceberg-python#4073 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
apache/iceberg-python#4010 · 3 commentaires · 1 réaction ·
Les mainteneurs répondent en général sous 1 jour
-
to_bytes silently rescales a Decimal with a negative scalePeut-être pris @Rodrigo-Palma l’a pris il y a 23 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3996 ·
Les mainteneurs répondent en général sous 1 jour
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationPeut-être pris @ghoshp83 l’a pris il y a 23 jours. Ouvertebug
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
apache/iceberg-python#3979 ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de apache/iceberg-python
Issues similaires
-
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
NousResearch/hermes-agent#136483 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 1/5 Moins d'une heure Accessibilité débutants 88/100
Les mainteneurs répondent en général sous 1 jour
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last elementPeut-être pris @peterdsharpe l’a pris aujourd’hui. Ouvertebug
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
pytorch/tensordict#2307 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
Les mainteneurs répondent en général sous 1 jour
-
GrokModel.generate/a_generate pass an OpenAI-style list-of-dicts to xai_sdk.chat.user(), so every call crashes with a protobuf TypeError before any network I/OPeut-être pris @Christian-Sidak l’a pris aujourd’hui. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
confident-ai/deepeval#3436 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour