Make manifest records schema-aware to avoid field mismatches across format versions
Mantenedores costumam responder em até 1 dia
@kevinjqliu já está trabalhando nisso.
Desde 12/9/2026.
Avaliação
Esta issue ainda não foi avaliada.
Descrição
Apache Iceberg version
main (development), commit 308768d99
Please describe the bug 🐞
Manifest records do not retain their schema, and their property getters use hard-coded field positions. When the in-memory record layout differs from the file schema, this can cause incorrect field access or silent data loss during serialization.
For example, referenced_data_file has field ID 143 in both v2 and v3, but its zero-based position differs:
| Field | v2 position | v3 position |
|---|---|---|
first_row_id |
Not present | 16 |
referenced_data_file |
16 | 17 |
This causes two problems:
- Writing a v3 record to a v2 file:
DataFile.from_args()defaults to the v3 layout. When writing it throughAvroOutputFilewith a v2 file schema, omittingrecord_schemamakes the writer read position 16 instead of 17 forreferenced_data_file. Iffirst_row_idis null, a non-null reference is silently written as null. Explicitly supplying the v3 record schema enables the existing field-ID projection and preserves the value. - Accessing a v2 record:
DataFile.from_args(_table_format_version=2, ...)stores the reference at position 16, but the getters still assume v3 positions. Consequently,first_row_idreturns the reference path andreferenced_data_fileraisesIndexError. Reading a v2 file intoDataFilewithout projecting to v3 has the same mismatch.
The normal manifest read/write helpers already supply the v3 projection. However, correctness still depends on callers separately tracking the in-memory layout and supplying the matching schema.
Proposed behavior
Records should keep the schema that describes their in-memory layout. Getters should use field IDs, not assume a particular version’s field positions.
When reading, use the schema from the file. If the read projects into another schema, retain that schema on the resulting record.
When writing, callers should only need to specify the target format version, or file_schema when using AvroOutputFile. IO should determine the input layout from the record and handle the conversion before encoding. This needs to account for nested records and partition fields, not just the top-level record.
We shouldn’t need to pass record_schema just to prevent a field from being read at the wrong position. Keep it supported for callers that explicitly provide it.
For the example above, writing a default v3 record to a v2 file should preserve referenced_data_file without passing record_schema. If a conversion isn’t supported, raise an error rather than silently writing the wrong value.
Willingness to contribute
N/A
- Linguagem predominante
- Python
- Estrelas
- 1.1k
- Forks
- 606
- Merge médio
- 1d 11h
- PRs com merge (30d)
- 75
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Tem um modelo de pull request
- Sem guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de apache/iceberg-python
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's responseTalvez já em andamento @Soumo-git-hub assumiu há 2 dias. Abertakind:bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 84/100
apache/iceberg-python#4073 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 70/100
apache/iceberg-python#4010 · 3 comentários · 1 reação ·
Mantenedores costumam responder em até 1 dia
-
to_bytes silently rescales a Decimal with a negative scaleTalvez já em andamento @Rodrigo-Palma assumiu há 18 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
apache/iceberg-python#3996 ·
Mantenedores costumam responder em até 1 dia
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationTalvez já em andamento @ghoshp83 assumiu há 18 dias. Abertabug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
apache/iceberg-python#3979 ·
Mantenedores costumam responder em até 1 dia
-
FsspecFileIO: `_adls` mutates shared properties, so a second storage account gets the first account's filesystemTalvez já em andamento @krishnakaanchan-png assumiu há 35 dias. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
apache/iceberg-python#3885 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de apache/iceberg-python
Issues semelhantes
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 92/100
QuantEcon/lecture-python-programming#642 ·
Mantenedores costumam responder em até 1 dia
-
area/config area/profiles comp/cli needs-decision P3 sweeper:risk-compatibility type/feature
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 82/100
NousResearch/hermes-agent#133697 ·
Mantenedores costumam responder em até 1 dia
-
enhancement needs-triage
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 72/100
Mantenedores costumam responder em até 1 dia
-
core
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 88/100
vectorize-io/hindsight#5279 ·
Mantenedores costumam responder em até 1 dia