Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

[Feature] Maintain primary-key full-text index archives on write and compaction

未關閉
#409 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
5/5
預估耗時
一週以上
新手友好度
25/100
Issue 類型
功能
描述清晰度
基本清楚
活躍度
活躍
技術堆疊
cpp
領域
databases

研究方向

Start with src/paimon/core/index/pk/bucketed_primary_key_index_maintainer.cpp and the level-planning code in src/paimon/core/index/pksorted/, after reviewing dependencies #400 and #408. Use the Java-aligned PkFullTextIndexFileTest, PkFullTextDataFileReaderTest, PkFullTextBucketIndexStateTest, and BucketedFullTextIndexMaintainerTest as behavioral references; done means archives are maintained through restore, commit, compaction, and abort, with the listed cross-language tests passing.

由索引模型根據 Issue 內容生成。

描述

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 5: primary-key full-text index, write side).

BucketedPrimaryKeyIndexMaintainer keeps only BTree definitions (src/paimon/core/index/pk/bucketed_primary_key_index_maintainer.cpp:128-132), and restore keeps only BTree payloads (IsPrimaryKeyBTreePayload). This has two effects:

  • Paimon C++ writers never build primary-key full-text archives.
  • When C++ compacts a table that already has Java-written archives, those archives are neither rebuilt nor retired. They keep pointing at data files that compaction has removed, and the new compacted files are not indexed.

Java model (apache/paimon#8649, apache/paimon#8651, apache/paimon#8672, apache/paimon#8992):

  • One archive per level. Each (partition, bucket, non-zero data level) has exactly one immutable archive. It covers all eligible files of that level, sorted by file name.
  • Eligible files. A file is eligible when fileSource == COMPACT && level > 0 (PrimaryKeyIndexSourcePolicy).
  • Row ids. An archive's row ids are the concatenated physical row positions of its source files, including rows deleted by deletion vectors and null rows.
  • PkFullTextIndexFile builds one archive:
    • index type full-text;
    • every source has the same level > 0 and a positive row count;
    • it calls GlobalIndexSingleColumnWriter#write(text, sourceOffset + rowPos) for each row;
    • finish() must return exactly one entry whose row count equals the total source row count.
    • The resulting IndexFileMeta has GlobalIndexMeta(0, total - 1, fieldId, null, indexMeta, sourceMeta):
      • indexMeta is the flat JSON of the prefix-stripped options;
      • sourceMeta is PrimaryKeyIndexSourceMeta(level, sourceFiles).
    • The file is named index-{uuid}-{N} under the index directory, or in the bucket directory when index-file-in-data-file-dir is set.
  • PkFullTextDataFileReader reads one text value per physical row, with no deletion-vector filtering. PkFullTextIndexBuilder builds an archive for one file or a list of files.
  • PkFullTextBucketIndexState#fromActiveDataFiles classifies payloads:
    • A payload is current only if its (level, ordered source files) exactly matches the level's eligible active files and its row counts match.
    • A payload that fails this is stale. So is every payload of a level that has more than one match, and any payload whose metadata cannot be parsed.
  • Level planning (PrimaryKeyIndexLevels, shared with the sorted indexes) picks the lowest level whose payload is missing or out of date. A plan with no source files removes the payload.
  • BucketedFullTextIndexMaintainer:
    • Restore: stale payloads are retired and emitted as deletions in the next commit.
    • prepareCommit(append, compact, waitCompaction):
      1. Applies the data transition: removes compactBefore files and adds eligible compactAfter files.
      2. Finishes or starts a single background build.
      3. Atomically replaces the level's archive when the build is still valid (canAccept); otherwise deletes the generated file.
      4. Routes the result to the compact increment if there was a compact transition, and to the append increment otherwise.
      5. Rolls back and deletes generated files on failure. The returned commit carries an abort hook.
    • A failed build is not retried within the same call. The next prepareCommit plans again.
  • Wiring:
    • BucketedPrimaryKeyIndexMaintainer.Factory creates the full-text maintainer for fixed-bucket writes and for postpone-bucket compaction (apache/paimon#8992).
    • IndexFileHandler#pkFullTextIndex(partition, bucket) provides the index file.
    • Writers are not closed while a build is pending.
Solution
  • Port the classes above.
  • Reuse the existing C++ PrimaryKeyIndexSourceMeta, PrimaryKeyIndexSourcePolicy and PrimaryKeyIndexSourceFile. Extract the level planning currently embedded in src/paimon/core/index/pksorted/ so it can be shared.
  • Wire the full-text maintainer into BucketedPrimaryKeyIndexMaintainer: restore, prepareCommit, abort, and merging the increments. Create it through the full-text indexer from #400, using the options resolved in #408.
  • Add tests aligned with Java PkFullTextIndexFileTest, PkFullTextDataFileReaderTest, PkFullTextBucketIndexStateTest and BucketedFullTextIndexMaintainerTest, plus:
    • C++ compaction of a table with Java-written primary-key full-text archives; stale archives are retired and new ones built;
    • Java reading archives written by C++.
Anything else?
  • Depends on #400 and #408.
  • The realtime path still rejects primary-key global indexes (src/paimon/core/utils/primary_key_table_utils.cpp:136-141). That is out of scope here.
Are you willing to submit a PR?
  • I'm willing to submit a PR!
主要語言
C++
星號
65
分支
31
平均合併
1 天 23 小時
30 天內合併 PR
64

環境準備

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

apache/paimon-cpp 的其他 Issue

查看 apache/paimon-cpp 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。