datalake_fdw: read path — planned fragments, CustomScan, projection and pruning
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
Línea de trabajo
Start with the engine plan and the CustomScan planner hook, using dependencies A, B2, and B4; validate the skeleton, projection, and pruning against the stub and local files first. Done means matching Spark rows across schema evolution and fragment splits, reporting correct EXPLAIN ANALYZE pruning, refusing unsupported or dictionary-encoded columns by name, and releasing readers on cancellation.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Part of #2008. Letters (A, B0–B7, C, D, E) are the PRs listed there; this is D.
Scope
- Fragments from the engine's plan (files, row-group ranges) assigned to segments; one large file may be split by row group.
- A CustomScan (planner hook, design decision D10) that opens a reader per fragment and decodes batches into slots.
- Projection by field id from the table's Iceberg schema (
ProjectionSet), so files written before a column was added or renamed read correctly; only needed columns are read. - Row-group pruning from the query's quals against Parquet statistics;
EXPLAIN ANALYZEreports row groups skipped. - With it: #1989 name mapping for files without field ids.
Out of scope
Delete files (E), time travel (#1683 §2.3), ANALYZE beyond the current zero-sample no-op.
Depends on
A, B2, B4. Skeleton, projection and pruning are testable against the stub and local files first.
Acceptance
- Same rows as Spark on the same table, including one Spark evolved (column added, renamed, dropped, int promoted to long).
- A file split into three fragments reads the same as whole; every segment reads only its fragments.
EXPLAIN ANALYZEshows pruning for a selective predicate and none forWHERE true.- Unsupported or dictionary-encoded columns are refused naming the column; cancel mid-scan releases the reader.
- Lenguaje dominante
- C
- Estrellas
- 1.4k
- Forks
- 248
- Merge medio
- 4 d 10 h
- PR fusionados (30 d)
- 40
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de apache/cloudberry
-
type: Bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
apache/cloudberry#1885 · 2 reacciones ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
apache/cloudberry#1825 ·
-
type: Bug
Dificultad 3/5 1-2 días Aptitud para principiantes 65/100
apache/cloudberry#2048 · 1 reacción ·
-
type: Bug
Dificultad 4/5 3-5 días Aptitud para principiantes 40/100
apache/cloudberry#2047 ·
-
type: Bug
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
apache/cloudberry#2046 · 1 comentario ·
Todos los issues de apache/cloudberry
Issues similares
-
task
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
vsanthanam/JBird#429 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
-
bug documentation
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
es-ude/OnDeviceTraining#459 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
bilelmoussaoui/gobject-linter#199 · 1 comentario ·
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
bradcypert/plum#53 ·