Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

[Feature] Support full-text search modes for unindexed row ranges

Ouverte
#406 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Les mainteneurs répondent en général sous 1 jour

Personne n'a encore pris cette issue.

Évaluation

Difficulté
5/5
Temps estimé
Plus d'une semaine
Accessibilité débutants
35/100
Type d'issue
Fonctionnalité
Clarté
Plutôt claire
Activité
Active
Stack technique
cpp
Domaine
databases, search

Piste de recherche

Start by reviewing the Java design references for DataEvolutionGlobalIndexCoverage and RawFullTextReadImpl, then inspect the C++ full-text search entry points after dependencies #404 and #405. Verify the four options and resolution order, raw-range handling, and in-memory temporary index behavior with tests for appended rows, partial coverage, deletion vectors, and row filters.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 4: search modes).

In Paimon C++, a full-text search only sees rows covered by a full-text index. Rows appended after the last index build are invisible. Java controls this with full-text-index.search-mode, which falls back to global-index.search-mode. In full and detail modes, Java also searches row ranges that have no full-text index yet: it reads the raw rows and builds a temporary index (apache/paimon#8316, apache/paimon#8844).

Java options (CoreOptions):

Key Default Description
global-index.search-mode (none) Fallback search mode for global index queries.
scalar-index.search-mode fast (apache/paimon#8891) Search mode for scalar index queries.
vector-index.search-mode fast Search mode for vector index queries.
full-text-index.search-mode fast Search mode for full-text index queries.

Values:

  • fast: search indexed data only.
  • full: use the snapshot's next row id and the global index coverage to find missing row ids, and scan raw data only when a gap exists.
  • detail: scan data files to find the exact unindexed rows.

Resolution order: an explicitly set family key wins, then global-index.search-mode, then the family default.

Java design:

  • Coverage. DataEvolutionGlobalIndexCoverage#unindexedRanges(fieldIds, ...) computes the gaps:

    • fast: none. The same applies when the snapshot's nextRowId is null or not positive.
    • full: [0, nextRowId - 1] minus the indexed ranges.
    • detail: the non-null row-id ranges of all data files (a ScanMode.ALL read that respects the partition filter), minus the indexed ranges.
    • Indexed ranges are intersected across the requested fields. Both index_field_id and extra_field_ids count as coverage.
  • Scan. When at least one full-text index file exists and the unindexed ranges are non-empty, the scan adds a RawFullTextSearchSplit(rowRanges). With no full-text index at all, the result is empty even in full mode.

  • Read. RawFullTextReadImpl handles the raw split:

    1. Read the column plus _ROW_ID for the raw ranges, pinned to the plan snapshot. The read respects deletion vectors, and the row filter is applied to build the include set.
    2. Build a temporary in-memory index with the column's index type and options, writing (text, rowId - first.from).
    3. Search it with the same query and limit.
    4. Replace indexed hits that fall inside the raw ranges with the raw hits.
    5. Apply the final top-k.

    The temporary index's statistics come from the raw rows only.

Solution
  • Add the four options and the family-specific resolution.
  • Port the coverage computation, RawFullTextSearchSplit, and the raw read path with the temporary index. The temporary index should be written to and read from memory, without touching table storage.
  • Use scalar-index.search-mode for row-filter coverage in #405.
  • Add tests:
    • rows appended after the index build, in fast, full and detail modes
    • partial index coverage across partitions
    • deletion vectors inside raw ranges
    • a row filter on the raw path
Anything else?
  • Depends on #404. The raw path supports row filters once #405 lands.
  • Java primary-key full-text search supports only fast and rejects the other modes; see #410.
Are you willing to submit a PR?
  • I'm willing to submit a PR!
Langage dominant
C++
Étoiles
65
Forks
31
Merge moyen
1 j 14 h
PR mergées (30 j)
60

Préparer son environnement

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de apache/paimon-cpp

Toutes les issues de apache/paimon-cpp

Issues similaires

Plus d'issues C++

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.