[Feature] Support full-text search modes for unindexed row ranges
維護者通常 1 天內回覆
還沒有人認領這個 Issue。
評估
研究方向
Start by reviewing the Java design references for DataEvolutionGlobalIndexCoverage and RawFullTextReadImpl, then inspect the C++ full-text search entry points after dependencies #404 and #405. Verify the four options and resolution order, raw-range handling, and in-memory temporary index behavior with tests for appended rows, partial coverage, deletion vectors, and row filters.
由索引模型根據 Issue 內容生成。
描述
Search before asking
- I searched in the issues and found nothing similar.
Motivation
Sub-issue of #399 (step 4: search modes).
In Paimon C++, a full-text search only sees rows covered by a full-text index. Rows appended after the last index build are invisible. Java controls this with full-text-index.search-mode, which falls back to global-index.search-mode. In full and detail modes, Java also searches row ranges that have no full-text index yet: it reads the raw rows and builds a temporary index (apache/paimon#8316, apache/paimon#8844).
Java options (CoreOptions):
| Key | Default | Description |
|---|---|---|
global-index.search-mode |
(none) | Fallback search mode for global index queries. |
scalar-index.search-mode |
fast (apache/paimon#8891) |
Search mode for scalar index queries. |
vector-index.search-mode |
fast |
Search mode for vector index queries. |
full-text-index.search-mode |
fast |
Search mode for full-text index queries. |
Values:
fast: search indexed data only.full: use the snapshot's next row id and the global index coverage to find missing row ids, and scan raw data only when a gap exists.detail: scan data files to find the exact unindexed rows.
Resolution order: an explicitly set family key wins, then global-index.search-mode, then the family default.
Java design:
-
Coverage.
DataEvolutionGlobalIndexCoverage#unindexedRanges(fieldIds, ...)computes the gaps:fast: none. The same applies when the snapshot'snextRowIdis null or not positive.full:[0, nextRowId - 1]minus the indexed ranges.detail: the non-null row-id ranges of all data files (aScanMode.ALLread that respects the partition filter), minus the indexed ranges.- Indexed ranges are intersected across the requested fields. Both
index_field_idandextra_field_idscount as coverage.
-
Scan. When at least one full-text index file exists and the unindexed ranges are non-empty, the scan adds a
RawFullTextSearchSplit(rowRanges). With no full-text index at all, the result is empty even infullmode. -
Read.
RawFullTextReadImplhandles the raw split:- Read the column plus
_ROW_IDfor the raw ranges, pinned to the plan snapshot. The read respects deletion vectors, and the row filter is applied to build the include set. - Build a temporary in-memory index with the column's index type and options, writing
(text, rowId - first.from). - Search it with the same query and limit.
- Replace indexed hits that fall inside the raw ranges with the raw hits.
- Apply the final top-k.
The temporary index's statistics come from the raw rows only.
- Read the column plus
Solution
- Add the four options and the family-specific resolution.
- Port the coverage computation,
RawFullTextSearchSplit, and the raw read path with the temporary index. The temporary index should be written to and read from memory, without touching table storage. - Use
scalar-index.search-modefor row-filter coverage in #405. - Add tests:
- rows appended after the index build, in
fast,fullanddetailmodes - partial index coverage across partitions
- deletion vectors inside raw ranges
- a row filter on the raw path
- rows appended after the index build, in
Anything else?
- Depends on #404. The raw path supports row filters once #405 lands.
- Java primary-key full-text search supports only
fastand rejects the other modes; see #410.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- 主要語言
- C++
- 星號
- 65
- 分支
- 31
- 平均合併
- 1 天 14 小時
- 30 天內合併 PR
- 60
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
apache/paimon-cpp 的其他 Issue
-
[Feature] Warm up next data file in ConcatBatchReader可能已有人在做 @SteNicholas 於 1 天前認領。 未關閉enhancement
apache/paimon-cpp#419 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer可能已有人在做 @SteNicholas 於 1 天前認領。 未關閉enhancement
apache/paimon-cpp#417 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
[Feature] Support writing MAP<K, BLOB> fields可能已有人在做 @SteNicholas 於 3 天前認領。 未關閉enhancement
apache/paimon-cpp#415 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
enhancement
難度 5/5 一週以上 新手友好度 35/100
apache/paimon-cpp#410 ·
維護者通常 1 天內回覆
-
enhancement
難度 5/5 一週以上 新手友好度 25/100
apache/paimon-cpp#409 ·
維護者通常 1 天內回覆
查看 apache/paimon-cpp 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 85/100
Icinga/icinga2#11077 · 1 則留言 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 66/100
-
agent:WSL bug linux LOW ui
難度 1/5 1 小時以內 新手友好度 78/100
維護者通常 1 天內回覆
-
Copter: PosHold brake-entry threshold became 16 deg instead of 0.16 deg after the radians conversion可能已有人在做 關聯的 PR 仍在進行中或已合併。 未關閉
難度 1/5 1 小時以內 新手友好度 78/100
ArduPilot/ardupilot#34617 · 1 則留言 · 1 個 reaction ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 85/100
tesseract-robotics/tesseract_nanobind#168 ·
維護者通常 1 天內回覆