datalake_fdw: read path — planned fragments, CustomScan, projection and pruning
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Domain
- data-engineering, databases, distributed-systems
Research direction
Start with the engine plan and the CustomScan planner hook, using dependencies A, B2, and B4; validate the skeleton, projection, and pruning against the stub and local files first. Done means matching Spark rows across schema evolution and fragment splits, reporting correct EXPLAIN ANALYZE pruning, refusing unsupported or dictionary-encoded columns by name, and releasing readers on cancellation.
Written by the indexing model from the issue text.
Description
Part of #2008. Letters (A, B0–B7, C, D, E) are the PRs listed there; this is D.
Scope
- Fragments from the engine's plan (files, row-group ranges) assigned to segments; one large file may be split by row group.
- A CustomScan (planner hook, design decision D10) that opens a reader per fragment and decodes batches into slots.
- Projection by field id from the table's Iceberg schema (
ProjectionSet), so files written before a column was added or renamed read correctly; only needed columns are read. - Row-group pruning from the query's quals against Parquet statistics;
EXPLAIN ANALYZEreports row groups skipped. - With it: #1989 name mapping for files without field ids.
Out of scope
Delete files (E), time travel (#1683 §2.3), ANALYZE beyond the current zero-sample no-op.
Depends on
A, B2, B4. Skeleton, projection and pruning are testable against the stub and local files first.
Acceptance
- Same rows as Spark on the same table, including one Spark evolved (column added, renamed, dropped, int promoted to long).
- A file split into three fragments reads the same as whole; every segment reads only its fragments.
EXPLAIN ANALYZEshows pruning for a selective predicate and none forWHERE true.- Unsupported or dictionary-encoded columns are refused naming the column; cancel mid-scan releases the reader.
- Dominant language
- C
- Stars
- 1.4k
- Forks
- 248
- Avg merge
- 4d 10h
- Merged PRs (30d)
- 40
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/cloudberry
-
type: Bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
apache/cloudberry#1885 · 2 reactions ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
apache/cloudberry#1825 ·
-
type: Bug
Difficulty 3/5 1-2 days Newbie friendliness 65/100
apache/cloudberry#2048 · 1 reaction ·
-
type: Bug
Difficulty 4/5 3-5 days Newbie friendliness 40/100
apache/cloudberry#2047 ·
-
type: Bug
Difficulty 4/5 3-5 days Newbie friendliness 45/100
apache/cloudberry#2046 · 1 comment ·
All issues in apache/cloudberry
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
bradcypert/plum#53 ·
-
Component: GLib
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Status: Opened
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
nextbsd/nextbsd-userland#285 ·