[Feature] Support full-text search modes for unindexed row ranges
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
调研方向
Start by reviewing the Java design references for DataEvolutionGlobalIndexCoverage and RawFullTextReadImpl, then inspect the C++ full-text search entry points after dependencies #404 and #405. Verify the four options and resolution order, raw-range handling, and in-memory temporary index behavior with tests for appended rows, partial coverage, deletion vectors, and row filters.
由索引模型根据 Issue 内容生成。
描述
Search before asking
- I searched in the issues and found nothing similar.
Motivation
Sub-issue of #399 (step 4: search modes).
In Paimon C++, a full-text search only sees rows covered by a full-text index. Rows appended after the last index build are invisible. Java controls this with full-text-index.search-mode, which falls back to global-index.search-mode. In full and detail modes, Java also searches row ranges that have no full-text index yet: it reads the raw rows and builds a temporary index (apache/paimon#8316, apache/paimon#8844).
Java options (CoreOptions):
| Key | Default | Description |
|---|---|---|
global-index.search-mode |
(none) | Fallback search mode for global index queries. |
scalar-index.search-mode |
fast (apache/paimon#8891) |
Search mode for scalar index queries. |
vector-index.search-mode |
fast |
Search mode for vector index queries. |
full-text-index.search-mode |
fast |
Search mode for full-text index queries. |
Values:
fast: search indexed data only.full: use the snapshot's next row id and the global index coverage to find missing row ids, and scan raw data only when a gap exists.detail: scan data files to find the exact unindexed rows.
Resolution order: an explicitly set family key wins, then global-index.search-mode, then the family default.
Java design:
-
Coverage.
DataEvolutionGlobalIndexCoverage#unindexedRanges(fieldIds, ...)computes the gaps:fast: none. The same applies when the snapshot'snextRowIdis null or not positive.full:[0, nextRowId - 1]minus the indexed ranges.detail: the non-null row-id ranges of all data files (aScanMode.ALLread that respects the partition filter), minus the indexed ranges.- Indexed ranges are intersected across the requested fields. Both
index_field_idandextra_field_idscount as coverage.
-
Scan. When at least one full-text index file exists and the unindexed ranges are non-empty, the scan adds a
RawFullTextSearchSplit(rowRanges). With no full-text index at all, the result is empty even infullmode. -
Read.
RawFullTextReadImplhandles the raw split:- Read the column plus
_ROW_IDfor the raw ranges, pinned to the plan snapshot. The read respects deletion vectors, and the row filter is applied to build the include set. - Build a temporary in-memory index with the column's index type and options, writing
(text, rowId - first.from). - Search it with the same query and limit.
- Replace indexed hits that fall inside the raw ranges with the raw hits.
- Apply the final top-k.
The temporary index's statistics come from the raw rows only.
- Read the column plus
Solution
- Add the four options and the family-specific resolution.
- Port the coverage computation,
RawFullTextSearchSplit, and the raw read path with the temporary index. The temporary index should be written to and read from memory, without touching table storage. - Use
scalar-index.search-modefor row-filter coverage in #405. - Add tests:
- rows appended after the index build, in
fast,fullanddetailmodes - partial index coverage across partitions
- deletion vectors inside raw ranges
- a row filter on the raw path
- rows appended after the index build, in
Anything else?
- Depends on #404. The raw path supports row filters once #405 lands.
- Java primary-key full-text search supports only
fastand rejects the other modes; see #410.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- 主要语言
- C++
- 星标
- 65
- 派生
- 31
- 平均合并
- 1 天 14 小时
- 30 天内合并 PR
- 60
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/paimon-cpp 的其他 Issue
-
[Feature] Warm up next data file in ConcatBatchReader可能已有人在做 @SteNicholas 于 1 天前认领。 未关闭enhancement
apache/paimon-cpp#419 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer可能已有人在做 @SteNicholas 于 1 天前认领。 未关闭enhancement
apache/paimon-cpp#417 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[Feature] Support writing MAP<K, BLOB> fields可能已有人在做 @SteNicholas 于 3 天前认领。 未关闭enhancement
apache/paimon-cpp#415 · 已指派 1 人 ·
维护者通常 1 天内回复
-
enhancement
难度 5/5 一周以上 新手友好度 35/100
apache/paimon-cpp#410 ·
维护者通常 1 天内回复
-
enhancement
难度 5/5 一周以上 新手友好度 25/100
apache/paimon-cpp#409 ·
维护者通常 1 天内回复
查看 apache/paimon-cpp 的全部 Issue
相似的 Issue
-
Make Catch2 optional when `RDK_BUILD_CPP_TESTS=OFF`可能已有人在做 @pechersky 今天认领。 未关闭bug
难度 2/5 1-3 小时 新手友好度 82/100
维护者通常 2 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 70/100
dice-group/dice-hash#111 ·
-
难度 2/5 1-3 小时 新手友好度 64/100
-
难度 2/5 1-3 小时 新手友好度 88/100
MerginMaps/mobile#4741 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
ros-perception/image_pipeline#1198 ·