Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Feature] Support full-text search modes for unindexed row ranges

オープン
#406 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
活発
技術スタック
cpp
領域
databases, search

調査の方向性

Start by reviewing the Java design references for DataEvolutionGlobalIndexCoverage and RawFullTextReadImpl, then inspect the C++ full-text search entry points after dependencies #404 and #405. Verify the four options and resolution order, raw-range handling, and in-memory temporary index behavior with tests for appended rows, partial coverage, deletion vectors, and row filters.

索引モデルが issue の本文から書いたものです。

説明

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 4: search modes).

In Paimon C++, a full-text search only sees rows covered by a full-text index. Rows appended after the last index build are invisible. Java controls this with full-text-index.search-mode, which falls back to global-index.search-mode. In full and detail modes, Java also searches row ranges that have no full-text index yet: it reads the raw rows and builds a temporary index (apache/paimon#8316, apache/paimon#8844).

Java options (CoreOptions):

Key Default Description
global-index.search-mode (none) Fallback search mode for global index queries.
scalar-index.search-mode fast (apache/paimon#8891) Search mode for scalar index queries.
vector-index.search-mode fast Search mode for vector index queries.
full-text-index.search-mode fast Search mode for full-text index queries.

Values:

  • fast: search indexed data only.
  • full: use the snapshot's next row id and the global index coverage to find missing row ids, and scan raw data only when a gap exists.
  • detail: scan data files to find the exact unindexed rows.

Resolution order: an explicitly set family key wins, then global-index.search-mode, then the family default.

Java design:

  • Coverage. DataEvolutionGlobalIndexCoverage#unindexedRanges(fieldIds, ...) computes the gaps:

    • fast: none. The same applies when the snapshot's nextRowId is null or not positive.
    • full: [0, nextRowId - 1] minus the indexed ranges.
    • detail: the non-null row-id ranges of all data files (a ScanMode.ALL read that respects the partition filter), minus the indexed ranges.
    • Indexed ranges are intersected across the requested fields. Both index_field_id and extra_field_ids count as coverage.
  • Scan. When at least one full-text index file exists and the unindexed ranges are non-empty, the scan adds a RawFullTextSearchSplit(rowRanges). With no full-text index at all, the result is empty even in full mode.

  • Read. RawFullTextReadImpl handles the raw split:

    1. Read the column plus _ROW_ID for the raw ranges, pinned to the plan snapshot. The read respects deletion vectors, and the row filter is applied to build the include set.
    2. Build a temporary in-memory index with the column's index type and options, writing (text, rowId - first.from).
    3. Search it with the same query and limit.
    4. Replace indexed hits that fall inside the raw ranges with the raw hits.
    5. Apply the final top-k.

    The temporary index's statistics come from the raw rows only.

Solution
  • Add the four options and the family-specific resolution.
  • Port the coverage computation, RawFullTextSearchSplit, and the raw read path with the temporary index. The temporary index should be written to and read from memory, without touching table storage.
  • Use scalar-index.search-mode for row-filter coverage in #405.
  • Add tests:
    • rows appended after the index build, in fast, full and detail modes
    • partial index coverage across partitions
    • deletion vectors inside raw ranges
    • a row filter on the raw path
Anything else?
  • Depends on #404. The raw path supports row filters once #405 lands.
  • Java primary-key full-text search supports only fast and rejects the other modes; see #410.
Are you willing to submit a PR?
  • I'm willing to submit a PR!
主要言語
C++
スター
66
フォーク
33
平均マージ
1日 14時間
マージ済み PR(30日)
57

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

apache/paimon-cpp のほかの issue

apache/paimon-cpp の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。