Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Tracking] datalake_fdw: Iceberg data path and the metadata agent

未关闭
#2,008 0 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

维护者通常 1 天内回复

@MisterRaindrop 已经在做这个了。

开始于 2026年9月16日。

评估

这个 Issue 还没有评估数据。

描述

datalake

Everything between the DDL skeleton (#1842) and a lake table that can be written and read from SQL. Design: #1683. This issue is the execution view.

Landed: #1842, the skeleton (AM, catalog/volume FDWs, stub engine, DDL guards); #1951, the Parquet format layer (Arrow, field ids, tracked memory pool, type rules).

Twelve PRs, three waves
 wave 1                           wave 2                            wave 3
 A  S3 storage ────────────────────────────────────────┐
 B0 proto ─┬► B1 agent framework ─┬► B3 Iceberg ops ─► B4 Builtin ─┼► C write ─┬► E merge-on-read + DML
           └► B2 engine client ───┘        ├► B5 Polaris            ├► D read ──┘
                                           ├► B6 Hive Metastore
                                           └► B7 Hadoop

C and D need A, B2 and B4; with the Builtin catalog they run in CI against MinIO and the agent alone. B5–B7 need B3 and run alongside C and D.

  • A #2009 Storage I/O over object storage (S3) — first
  • B0 #2010 The gRPC contract: proto files
  • B1 #2011 datalake_agent: service framework (Java)
  • B2 #2012 datalake_fdw: the agent engine client
  • B3 #2013 datalake_agent: Iceberg operations by metadata location
  • B4 #2014 datalake_fdw: Builtin catalog
  • B5 #2015 datalake_agent: Polaris (REST) catalog
  • B6 #2016 datalake_agent: Hive Metastore catalog
  • B7 #2017 datalake_agent: Hadoop (filesystem) catalog
  • C #2018 Write path: INSERT to data files and a committed snapshot
  • D #2019 Read path: planned fragments, CustomScan, projection and pruning
  • E #2020 Merge-on-read and DML

Attached to the PR they belong with: #1988 NUMERIC (C), #1990 timestamp units (C), #1989 name mapping (D). Later, blocking nothing: the optional datalake_proxy bgworker (#1683 §5.4).

Rules
  • PRs target main and are squash-merged; the description is the commit message (#1951 layout).
  • Every PR leaves both components shippable: an unfinished path refuses with a clear error. The agent answers UNIMPLEMENTED for RPCs it does not carry yet, as the stub engine does.
  • The proto files and the shared headers (format/format.h, meta/iceberg_meta_engine.h, common/*.h) change only in their own small PR.
  • Nothing generated is committed; protoc runs at build time on both sides.
  • At most two PRs in review at once; the next stays a draft.
  • Each PR brings its tests: a smoke category on the C side, mvn -B verify on the Java side, a CI job for any service it needs.
  • Milestone datalake_fdw: Iceberg read/write MVP, label datalake.

Out of scope: the FDW raw-file path (#1683 §2.3), the CI build image (devops repo), HDFS storage, execution-engine work.

主要语言
C
星标
1.4k
派生
248
平均合并
4 天 13 小时
30 天内合并 PR
38

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/cloudberry 的其他 Issue

查看 apache/cloudberry 的全部 Issue

相似的 Issue

更多 C Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。