Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

[Feature] Support full-text search modes for unindexed row ranges

未關閉
#406 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
5/5
預估耗時
一週以上
新手友好度
35/100
Issue 類型
功能
描述清晰度
基本清楚
活躍度
活躍
技術堆疊
cpp
領域
databases, search

研究方向

Start by reviewing the Java design references for DataEvolutionGlobalIndexCoverage and RawFullTextReadImpl, then inspect the C++ full-text search entry points after dependencies #404 and #405. Verify the four options and resolution order, raw-range handling, and in-memory temporary index behavior with tests for appended rows, partial coverage, deletion vectors, and row filters.

由索引模型根據 Issue 內容生成。

描述

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 4: search modes).

In Paimon C++, a full-text search only sees rows covered by a full-text index. Rows appended after the last index build are invisible. Java controls this with full-text-index.search-mode, which falls back to global-index.search-mode. In full and detail modes, Java also searches row ranges that have no full-text index yet: it reads the raw rows and builds a temporary index (apache/paimon#8316, apache/paimon#8844).

Java options (CoreOptions):

Key Default Description
global-index.search-mode (none) Fallback search mode for global index queries.
scalar-index.search-mode fast (apache/paimon#8891) Search mode for scalar index queries.
vector-index.search-mode fast Search mode for vector index queries.
full-text-index.search-mode fast Search mode for full-text index queries.

Values:

  • fast: search indexed data only.
  • full: use the snapshot's next row id and the global index coverage to find missing row ids, and scan raw data only when a gap exists.
  • detail: scan data files to find the exact unindexed rows.

Resolution order: an explicitly set family key wins, then global-index.search-mode, then the family default.

Java design:

  • Coverage. DataEvolutionGlobalIndexCoverage#unindexedRanges(fieldIds, ...) computes the gaps:

    • fast: none. The same applies when the snapshot's nextRowId is null or not positive.
    • full: [0, nextRowId - 1] minus the indexed ranges.
    • detail: the non-null row-id ranges of all data files (a ScanMode.ALL read that respects the partition filter), minus the indexed ranges.
    • Indexed ranges are intersected across the requested fields. Both index_field_id and extra_field_ids count as coverage.
  • Scan. When at least one full-text index file exists and the unindexed ranges are non-empty, the scan adds a RawFullTextSearchSplit(rowRanges). With no full-text index at all, the result is empty even in full mode.

  • Read. RawFullTextReadImpl handles the raw split:

    1. Read the column plus _ROW_ID for the raw ranges, pinned to the plan snapshot. The read respects deletion vectors, and the row filter is applied to build the include set.
    2. Build a temporary in-memory index with the column's index type and options, writing (text, rowId - first.from).
    3. Search it with the same query and limit.
    4. Replace indexed hits that fall inside the raw ranges with the raw hits.
    5. Apply the final top-k.

    The temporary index's statistics come from the raw rows only.

Solution
  • Add the four options and the family-specific resolution.
  • Port the coverage computation, RawFullTextSearchSplit, and the raw read path with the temporary index. The temporary index should be written to and read from memory, without touching table storage.
  • Use scalar-index.search-mode for row-filter coverage in #405.
  • Add tests:
    • rows appended after the index build, in fast, full and detail modes
    • partial index coverage across partitions
    • deletion vectors inside raw ranges
    • a row filter on the raw path
Anything else?
  • Depends on #404. The raw path supports row filters once #405 lands.
  • Java primary-key full-text search supports only fast and rejects the other modes; see #410.
Are you willing to submit a PR?
  • I'm willing to submit a PR!
主要語言
C++
星號
65
分支
31
平均合併
1 天 14 小時
30 天內合併 PR
60

環境準備

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

apache/paimon-cpp 的其他 Issue

查看 apache/paimon-cpp 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。