datalake_fdw: read path — planned fragments, CustomScan, projection and pruning
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Ambito
- data-engineering, databases, distributed-systems
Direzione di ricerca
Start with the engine plan and the CustomScan planner hook, using dependencies A, B2, and B4; validate the skeleton, projection, and pruning against the stub and local files first. Done means matching Spark rows across schema evolution and fragment splits, reporting correct EXPLAIN ANALYZE pruning, refusing unsupported or dictionary-encoded columns by name, and releasing readers on cancellation.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Part of #2008. Letters (A, B0–B7, C, D, E) are the PRs listed there; this is D.
Scope
- Fragments from the engine's plan (files, row-group ranges) assigned to segments; one large file may be split by row group.
- A CustomScan (planner hook, design decision D10) that opens a reader per fragment and decodes batches into slots.
- Projection by field id from the table's Iceberg schema (
ProjectionSet), so files written before a column was added or renamed read correctly; only needed columns are read. - Row-group pruning from the query's quals against Parquet statistics;
EXPLAIN ANALYZEreports row groups skipped. - With it: #1989 name mapping for files without field ids.
Out of scope
Delete files (E), time travel (#1683 §2.3), ANALYZE beyond the current zero-sample no-op.
Depends on
A, B2, B4. Skeleton, projection and pruning are testable against the stub and local files first.
Acceptance
- Same rows as Spark on the same table, including one Spark evolved (column added, renamed, dropped, int promoted to long).
- A file split into three fragments reads the same as whole; every segment reads only its fragments.
EXPLAIN ANALYZEshows pruning for a selective predicate and none forWHERE true.- Unsupported or dictionary-encoded columns are refused naming the column; cancel mid-scan releases the reader.
- Lingua principale
- C
- Stelle
- 1.4k
- Fork
- 248
- Merge medio
- 4g 10h
- PR unite (30g)
- 40
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di apache/cloudberry
-
type: Bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
apache/cloudberry#1885 · 2 reazioni ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
apache/cloudberry#1825 ·
-
type: Bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 65/100
apache/cloudberry#2048 · 1 reazione ·
-
type: Bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 40/100
apache/cloudberry#2047 ·
-
type: Bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
apache/cloudberry#2046 · 1 commento ·
Tutte le issue di apache/cloudberry
Issue simili
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
bradcypert/plum#53 ·
-
Component: GLib
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Status: Opened
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
nextbsd/nextbsd-userland#285 ·