[Feature] Maintain source-backed primary-key BTree indexes during compaction

未关闭
#291 0 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
25/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
停滞
技术栈
cpp
领域
databases

调研方向

首先阅读 #245 中引用的实现,以及 #192 和 #194 中由源代码支持的 primary-key read path。跟踪现有的存储格式,以及内部 reader 和 writer、manifest、path、sort-buffer 和 commit 抽象。完成标准是 fixed-bucket writes 和 compaction 能够维护按字段划分且位于正层级的 payload,同时 cleanup 和失败的 builds 保留安全的 scan fallback。

由索引模型根据 Issue 内容生成。

描述

Search before asking
  • I searched in the issues and found nothing similar.
Motivation

#192 and #194 added the source-backed primary-key BTree read path. Paimon C++ writers still need the corresponding maintenance path: after compaction changes the active source files of a data level, a missing or stale payload leaves that level uncovered and queries fall back to normal file scans.

Paimon C++ should maintain these payloads during fixed-bucket primary-key writes and compaction, using the existing Java-compatible source metadata, BTree payload format, and index manifests.

Solution

Add the source-backed primary-key BTree maintenance lifecycle for fixed-bucket primary-key tables:

  1. Validate the Java-equivalent table and index prerequisites.
  2. Restore committed source-backed payload metadata into bucket writers without mixing Data Evolution payloads.
  3. Build one payload per indexed field and positive data level from physical source rows, then commit matching index additions and deletions in the same snapshot as the data changes.
  4. Reconcile missing, stale, duplicate, replaced, removed-definition, and empty-level payloads during compaction.
  5. Isolate build failures to the affected field and level so reads safely fall back to normal scans and a later maintenance attempt can rebuild the payload.
  6. Retain live index files during snapshot expiration and orphan cleanup, including tag and branch safety and external-file deletion retries.

Reuse the existing storage formats and internal reader, writer, sort-buffer, path, manifest, and commit abstractions. Do not introduce a new index family or storage protocol.

The implementation is in #245. It keeps maintenance synchronous; Java asynchronous scheduling, manual rebuild actions, realtime writers, and postpone-bucket writers remain outside this scope.

Anything else?

This is a maintenance-path follow-up to the read-path work in #192 and #194. It ports an existing Java capability, so no separate PIP is proposed.

Are you willing to submit a PR?
  • I'm willing to submit a PR!
主要语言
C++
星标
65
派生
29
平均合并
2 天 30 分钟
30 天内合并 PR
77

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/paimon-cpp 的其他 Issue

查看 apache/paimon-cpp 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。