Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

datalake_fdw: read path — planned fragments, CustomScan, projection and pruning

Open
#2,019 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
c, postgresql, sql

Research direction

Start with the engine plan and the CustomScan planner hook, using dependencies A, B2, and B4; validate the skeleton, projection, and pruning against the stub and local files first. Done means matching Spark rows across schema evolution and fragment splits, reporting correct EXPLAIN ANALYZE pruning, refusing unsupported or dictionary-encoded columns by name, and releasing readers on cancellation.

Written by the indexing model from the issue text.

Description

datalake

Part of #2008. Letters (A, B0–B7, C, D, E) are the PRs listed there; this is D.

Scope
  • Fragments from the engine's plan (files, row-group ranges) assigned to segments; one large file may be split by row group.
  • A CustomScan (planner hook, design decision D10) that opens a reader per fragment and decodes batches into slots.
  • Projection by field id from the table's Iceberg schema (ProjectionSet), so files written before a column was added or renamed read correctly; only needed columns are read.
  • Row-group pruning from the query's quals against Parquet statistics; EXPLAIN ANALYZE reports row groups skipped.
  • With it: #1989 name mapping for files without field ids.
Out of scope

Delete files (E), time travel (#1683 §2.3), ANALYZE beyond the current zero-sample no-op.

Depends on

A, B2, B4. Skeleton, projection and pruning are testable against the stub and local files first.

Acceptance
  • Same rows as Spark on the same table, including one Spark evolved (column added, renamed, dropped, int promoted to long).
  • A file split into three fragments reads the same as whole; every segment reads only its fragments.
  • EXPLAIN ANALYZE shows pruning for a selective predicate and none for WHERE true.
  • Unsupported or dictionary-encoded columns are refused naming the column; cancel mid-scan releases the reader.
Dominant language
C
Stars
1.4k
Forks
248
Avg merge
4d 10h
Merged PRs (30d)
40

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/cloudberry

All issues in apache/cloudberry

Similar issues

More C issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.