Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Feature] Maintain primary-key full-text index archives on write and compaction

未关闭
#409 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
25/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
活跃
技术栈
cpp
领域
databases

调研方向

Start with src/paimon/core/index/pk/bucketed_primary_key_index_maintainer.cpp and the level-planning code in src/paimon/core/index/pksorted/, after reviewing dependencies #400 and #408. Use the Java-aligned PkFullTextIndexFileTest, PkFullTextDataFileReaderTest, PkFullTextBucketIndexStateTest, and BucketedFullTextIndexMaintainerTest as behavioral references; done means archives are maintained through restore, commit, compaction, and abort, with the listed cross-language tests passing.

由索引模型根据 Issue 内容生成。

描述

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 5: primary-key full-text index, write side).

BucketedPrimaryKeyIndexMaintainer keeps only BTree definitions (src/paimon/core/index/pk/bucketed_primary_key_index_maintainer.cpp:128-132), and restore keeps only BTree payloads (IsPrimaryKeyBTreePayload). This has two effects:

  • Paimon C++ writers never build primary-key full-text archives.
  • When C++ compacts a table that already has Java-written archives, those archives are neither rebuilt nor retired. They keep pointing at data files that compaction has removed, and the new compacted files are not indexed.

Java model (apache/paimon#8649, apache/paimon#8651, apache/paimon#8672, apache/paimon#8992):

  • One archive per level. Each (partition, bucket, non-zero data level) has exactly one immutable archive. It covers all eligible files of that level, sorted by file name.
  • Eligible files. A file is eligible when fileSource == COMPACT && level > 0 (PrimaryKeyIndexSourcePolicy).
  • Row ids. An archive's row ids are the concatenated physical row positions of its source files, including rows deleted by deletion vectors and null rows.
  • PkFullTextIndexFile builds one archive:
    • index type full-text;
    • every source has the same level > 0 and a positive row count;
    • it calls GlobalIndexSingleColumnWriter#write(text, sourceOffset + rowPos) for each row;
    • finish() must return exactly one entry whose row count equals the total source row count.
    • The resulting IndexFileMeta has GlobalIndexMeta(0, total - 1, fieldId, null, indexMeta, sourceMeta):
      • indexMeta is the flat JSON of the prefix-stripped options;
      • sourceMeta is PrimaryKeyIndexSourceMeta(level, sourceFiles).
    • The file is named index-{uuid}-{N} under the index directory, or in the bucket directory when index-file-in-data-file-dir is set.
  • PkFullTextDataFileReader reads one text value per physical row, with no deletion-vector filtering. PkFullTextIndexBuilder builds an archive for one file or a list of files.
  • PkFullTextBucketIndexState#fromActiveDataFiles classifies payloads:
    • A payload is current only if its (level, ordered source files) exactly matches the level's eligible active files and its row counts match.
    • A payload that fails this is stale. So is every payload of a level that has more than one match, and any payload whose metadata cannot be parsed.
  • Level planning (PrimaryKeyIndexLevels, shared with the sorted indexes) picks the lowest level whose payload is missing or out of date. A plan with no source files removes the payload.
  • BucketedFullTextIndexMaintainer:
    • Restore: stale payloads are retired and emitted as deletions in the next commit.
    • prepareCommit(append, compact, waitCompaction):
      1. Applies the data transition: removes compactBefore files and adds eligible compactAfter files.
      2. Finishes or starts a single background build.
      3. Atomically replaces the level's archive when the build is still valid (canAccept); otherwise deletes the generated file.
      4. Routes the result to the compact increment if there was a compact transition, and to the append increment otherwise.
      5. Rolls back and deletes generated files on failure. The returned commit carries an abort hook.
    • A failed build is not retried within the same call. The next prepareCommit plans again.
  • Wiring:
    • BucketedPrimaryKeyIndexMaintainer.Factory creates the full-text maintainer for fixed-bucket writes and for postpone-bucket compaction (apache/paimon#8992).
    • IndexFileHandler#pkFullTextIndex(partition, bucket) provides the index file.
    • Writers are not closed while a build is pending.
Solution
  • Port the classes above.
  • Reuse the existing C++ PrimaryKeyIndexSourceMeta, PrimaryKeyIndexSourcePolicy and PrimaryKeyIndexSourceFile. Extract the level planning currently embedded in src/paimon/core/index/pksorted/ so it can be shared.
  • Wire the full-text maintainer into BucketedPrimaryKeyIndexMaintainer: restore, prepareCommit, abort, and merging the increments. Create it through the full-text indexer from #400, using the options resolved in #408.
  • Add tests aligned with Java PkFullTextIndexFileTest, PkFullTextDataFileReaderTest, PkFullTextBucketIndexStateTest and BucketedFullTextIndexMaintainerTest, plus:
    • C++ compaction of a table with Java-written primary-key full-text archives; stale archives are retired and new ones built;
    • Java reading archives written by C++.
Anything else?
  • Depends on #400 and #408.
  • The realtime path still rejects primary-key global indexes (src/paimon/core/utils/primary_key_table_utils.cpp:136-141). That is out of scope here.
Are you willing to submit a PR?
  • I'm willing to submit a PR!
主要语言
C++
星标
65
派生
31
平均合并
1 天 23 小时
30 天内合并 PR
64

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/paimon-cpp 的其他 Issue

查看 apache/paimon-cpp 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。