Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

Make manifest records schema-aware to avoid field mismatches across format versions

Offen
#3,957 0 Kommentare 0 Reaktionen 1 zugewiesene Person Auf GitHub ansehen

Maintainer antworten meist innerhalb von 1 Tag

@kevinjqliu arbeitet bereits daran.

Seit 12.9.2026.

Bewertung

Dieses Issue wurde noch nicht bewertet.

Beschreibung

bug
Apache Iceberg version

main (development), commit 308768d99

Please describe the bug 🐞

Manifest records do not retain their schema, and their property getters use hard-coded field positions. When the in-memory record layout differs from the file schema, this can cause incorrect field access or silent data loss during serialization.

For example, referenced_data_file has field ID 143 in both v2 and v3, but its zero-based position differs:

Field v2 position v3 position
first_row_id Not present 16
referenced_data_file 16 17

This causes two problems:

  • Writing a v3 record to a v2 file: DataFile.from_args() defaults to the v3 layout. When writing it through AvroOutputFile with a v2 file schema, omitting record_schema makes the writer read position 16 instead of 17 for referenced_data_file. If first_row_id is null, a non-null reference is silently written as null. Explicitly supplying the v3 record schema enables the existing field-ID projection and preserves the value.
  • Accessing a v2 record: DataFile.from_args(_table_format_version=2, ...) stores the reference at position 16, but the getters still assume v3 positions. Consequently, first_row_id returns the reference path and referenced_data_file raises IndexError. Reading a v2 file into DataFile without projecting to v3 has the same mismatch.

The normal manifest read/write helpers already supply the v3 projection. However, correctness still depends on callers separately tracking the in-memory layout and supplying the matching schema.

Proposed behavior

Records should keep the schema that describes their in-memory layout. Getters should use field IDs, not assume a particular version’s field positions.

When reading, use the schema from the file. If the read projects into another schema, retain that schema on the resulting record.

When writing, callers should only need to specify the target format version, or file_schema when using AvroOutputFile. IO should determine the input layout from the record and handle the conversion before encoding. This needs to account for nested records and partition fields, not just the top-level record.

We shouldn’t need to pass record_schema just to prevent a field from being read at the wrong position. Keep it supported for callers that explicitly provide it.

For the example above, writing a default v3 record to a v2 file should preserve referenced_data_file without passing record_schema. If a conversion isn’t supported, raise an error rather than silently writing the wrong value.

Willingness to contribute

N/A

Vorherrschende Sprache
Python
Sterne
1.1k
Forks
606
Ø Merge
1 T. 11 Std.
Gemergte PRs (30 T.)
76

Entwicklungsumgebung

  • Kein Dockerfile und keine Docker-Compose-Datei
  • Hat eine Pull-Request-Vorlage
  • Kein Beitragsleitfaden

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus apache/iceberg-python

Alle Issues in apache/iceberg-python

Ähnliche Issues

Weitere Issues zu Python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.