datalake_fdw: read path — planned fragments, CustomScan, projection and pruning
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Lĩnh vực
- data-engineering, databases, distributed-systems
Hướng nghiên cứu
Start with the engine plan and the CustomScan planner hook, using dependencies A, B2, and B4; validate the skeleton, projection, and pruning against the stub and local files first. Done means matching Spark rows across schema evolution and fragment splits, reporting correct EXPLAIN ANALYZE pruning, refusing unsupported or dictionary-encoded columns by name, and releasing readers on cancellation.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Part of #2008. Letters (A, B0–B7, C, D, E) are the PRs listed there; this is D.
Scope
- Fragments from the engine's plan (files, row-group ranges) assigned to segments; one large file may be split by row group.
- A CustomScan (planner hook, design decision D10) that opens a reader per fragment and decodes batches into slots.
- Projection by field id from the table's Iceberg schema (
ProjectionSet), so files written before a column was added or renamed read correctly; only needed columns are read. - Row-group pruning from the query's quals against Parquet statistics;
EXPLAIN ANALYZEreports row groups skipped. - With it: #1989 name mapping for files without field ids.
Out of scope
Delete files (E), time travel (#1683 §2.3), ANALYZE beyond the current zero-sample no-op.
Depends on
A, B2, B4. Skeleton, projection and pruning are testable against the stub and local files first.
Acceptance
- Same rows as Spark on the same table, including one Spark evolved (column added, renamed, dropped, int promoted to long).
- A file split into three fragments reads the same as whole; every segment reads only its fragments.
EXPLAIN ANALYZEshows pruning for a selective predicate and none forWHERE true.- Unsupported or dictionary-encoded columns are refused naming the column; cancel mid-scan releases the reader.
- Ngôn ngữ chính
- C
- Star
- 1.4k
- Fork
- 248
- Merge trung bình
- 4 ngày 10 giờ
- Pull request đã merge (30 ngày)
- 40
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/cloudberry
-
type: Bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
apache/cloudberry#1885 · 2 reaction ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
apache/cloudberry#1825 ·
-
type: Bug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 65/100
apache/cloudberry#2048 · 1 reaction ·
-
type: Bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 40/100
apache/cloudberry#2047 ·
-
type: Bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
apache/cloudberry#2046 · 1 bình luận ·
Tất cả issue của apache/cloudberry
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
bradcypert/plum#53 ·
-
Component: GLib
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
Status: Opened
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
nextbsd/nextbsd-userland#285 ·