datalake_fdw: Iceberg name mapping for data files without field ids
@MisterRaindrop ci sta già lavorando.
Dal 13/9/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Summary
Since #1951 the Parquet reader in contrib/datalake_fdw matches a table's columns to a file's by Iceberg field id (PARQUET:field_id in the Parquet schema). A file whose columns carry no field id can never be matched: every projected column reads as NULL. That is correct per spec for a file the table never wrote, and wrong for the one case the spec covers: tables created over existing Parquet data (add_files, migrated Hive tables), whose files predate the ids.
What the spec says
Iceberg resolves such files through the table property schema.name-mapping.default: a JSON mapping from field ids to the column names to look for in the file, including nested names and multiple names per id (for renamed columns). A reader applies it only to columns that have no field id.
What has to happen
- The metadata engine has to surface the property to the access method.
ProjectionSet(format/format.h) needs a way to hand the reader a name mapping alongside the field ids -- a second array of names per id, or a pointer to a parsed mapping.parquet_read.cpp'sparquet_project()matches by id first and falls back to the mapping for id-less columns. Nothing changes in the decoder.- A file with neither ids nor a mapping match stays NULL, as now.
Deferred from #1951 on purpose: the framework there is settled, and the property needs the metadata engine, which is a later PR.
- Lingua principale
- C
- Stelle
- 1.4k
- Fork
- 248
- Merge medio
- 4g 10h
- PR unite (30g)
- 40
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di apache/cloudberry
-
type: Bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
apache/cloudberry#1885 · 2 reazioni ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
apache/cloudberry#1825 ·
-
type: Bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 65/100
apache/cloudberry#2048 · 1 reazione ·
-
type: Bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 40/100
apache/cloudberry#2047 ·
-
type: Bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
apache/cloudberry#2046 · 1 commento ·
Tutte le issue di apache/cloudberry
Issue simili
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
bradcypert/plum#53 ·
-
Component: GLib
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Status: Opened
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
nextbsd/nextbsd-userland#285 ·