[Feature] Search primary-key full-text indexes
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
调研方向
Start with the rejection in src/paimon/core/table/source/primary_key_sorted_index_scan.cpp and the TODO in src/paimon/core/operation/raw_file_split_read.cpp, then review dependencies #404 and #409. Use the Java-aligned PrimaryKeyFullTextScanTest, PrimaryKeyFullTextReadTest, PrimaryKeyFullTextSearchTest, PrimaryKeyFullTextBucketSearchTest, and NativePrimaryKeyFullTextIndexTest as behavioral guides. Done means primary-key full-text searches plan, rank, read, and propagate scores with the listed filter, mode, deletion, bucket, and archive cases covered.
由索引模型根据 Issue 内容生成。
描述
Search before asking
- I searched in the issues and found nothing similar.
Motivation
Sub-issue of #399 (step 5: primary-key full-text index, read side).
Paimon C++ cannot search primary-key full-text indexes:
PrimaryKeySortedIndexScanrejects full-text search (src/paimon/core/table/source/primary_key_sorted_index_scan.cpp:288-290).- There is no scan, split, bucket search or read for
full-textpayloads. - Indexed scores are not propagated through the primary-key physical-position read path (the TODO in
src/paimon/core/operation/raw_file_split_read.cpp:85-89).
Java design (apache/paimon#8649, apache/paimon#8652, apache/paimon#8659, apache/paimon#8844, apache/paimon#9060, apache/paimon#9184):
- Dispatch.
FullTextSearchBuilderImplroutes to the primary-key path when the table is not a data-evolution table andpk-full-text.index.columnscovers the column. A non-partition filter on that path is rejected withPrimary-key full-text search does not support non-partition filters yet. PrimaryKeyFullTextScan:- Plan the primary-key batch scan with the partition filter, pinned to the snapshot.
- Scan the index manifest for
full-textpayloads that have source metadata and the definition's field id. All entries must beADD. - Group data splits by (partition, bucket), skipping
bucket < 0. - Keep eligible files with their aligned
DeletionFiles, and resolve current payloads withPkFullTextBucketIndexState#fromActiveDataFiles. - Emit one split per bucket.
PrimaryKeyFullTextSearchSplitholds the data split, the payload files and the uncovered data file names. Every eligible file is either covered by exactly one payload or listed as uncovered.PrimaryKeyFullTextBucketSearch(searchRankingsAsync):- For each payload, lay out its source files in order to get their row offsets.
- If a source is inactive or has deletions, the include set is the active ranges minus deleted positions. A payload whose include set is empty is skipped.
- Call
visitFullTextSearch(new FullTextSearch(column, query, limit).withIncludeRowIds(include)). - Map hits to
PrimaryKeySearchPosition(partition, bucket, fileName, rowId - offset, score), sorted by score descending, then file name, then position.
PrimaryKeyFullTextRead:- Requires
limit > 0, and supports onlyfull-text-index.search-mode = fast;full/detailthrowUnsupportedOperationException. - Searches each split asynchronously on the
global-index.thread-numexecutor, submitting from the caller thread. - Takes the global top-k with
PrimaryKeySearchRanker#topKByScore. - Returns
PrimaryKeyScoredResult. It turns into one indexed split per data file, with row ranges, scores and that file'sDeletionFile, which are read by position. - Uncovered data files are not searched in
fastmode.
- Requires
Solution
- Port
PrimaryKeyFullTextScan,PrimaryKeyFullTextSearchSplit(serializable),PrimaryKeyFullTextBucketSearchandPrimaryKeyFullTextRead. - Port the shared primitives
PrimaryKeySearchPosition,PrimaryKeySearchRanker#topKByScoreandPrimaryKeyScoredResult. Hybrid search (#407) will reuse them. - Propagate
_INDEX_SCOREthrough the primary-key positional read path. - Add primary-key dispatch to the table-level full-text builder from #404.
- Add tests aligned with Java
PrimaryKeyFullTextScanTest,PrimaryKeyFullTextReadTest,PrimaryKeyFullTextSearchTest,PrimaryKeyFullTextBucketSearchTestandNativePrimaryKeyFullTextIndexTest:- deletion vectors
- several buckets and levels
- uncovered files
- global top-k
- partition filters
- rejection of non-partition filters and non-
fastmodes - archives written by Java
Anything else?
Depends on #404 and #409.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- 主要语言
- C++
- 星标
- 65
- 派生
- 31
- 平均合并
- 1 天 14 小时
- 30 天内合并 PR
- 60
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/paimon-cpp 的其他 Issue
-
[Feature] Warm up next data file in ConcatBatchReader可能已有人在做 @SteNicholas 于 2 天前认领。 未关闭enhancement
apache/paimon-cpp#419 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer可能已有人在做 @SteNicholas 于 2 天前认领。 未关闭enhancement
apache/paimon-cpp#417 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[Feature] Support writing MAP<K, BLOB> fields可能已有人在做 @SteNicholas 于 4 天前认领。 未关闭enhancement
apache/paimon-cpp#415 · 已指派 1 人 ·
维护者通常 1 天内回复
-
enhancement
难度 5/5 一周以上 新手友好度 25/100
apache/paimon-cpp#409 ·
维护者通常 1 天内回复
-
enhancement
难度 5/5 一周以上 新手友好度 35/100
apache/paimon-cpp#408 ·
维护者通常 1 天内回复
查看 apache/paimon-cpp 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 70/100
shadps4-emu/shadps4-qtlauncher#465 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 78/100
Qiskit/qiskit-aer#2466 ·
-
bug needs-triage
难度 2/5 1-3 小时 新手友好度 75/100
NVIDIA/attestation-sdk#41 · 1 条评论 ·
-
[Bug]: CMAKE Fails to find libgit2 on POP OS可能已有人在做 @Tirpitz93 今天认领。 未关闭bug
难度 1/5 1 小时以内 新手友好度 85/100
subsurface/subsurface#5007 · 2 条评论 ·
维护者通常 2 天内回复
-
feature request
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 2 天内回复