Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[Feature] Search primary-key full-text indexes

Aperta
#410 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
35/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
cpp

Direzione di ricerca

Start with the rejection in src/paimon/core/table/source/primary_key_sorted_index_scan.cpp and the TODO in src/paimon/core/operation/raw_file_split_read.cpp, then review dependencies #404 and #409. Use the Java-aligned PrimaryKeyFullTextScanTest, PrimaryKeyFullTextReadTest, PrimaryKeyFullTextSearchTest, PrimaryKeyFullTextBucketSearchTest, and NativePrimaryKeyFullTextIndexTest as behavioral guides. Done means primary-key full-text searches plan, rank, read, and propagate scores with the listed filter, mode, deletion, bucket, and archive cases covered.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 5: primary-key full-text index, read side).

Paimon C++ cannot search primary-key full-text indexes:

  • PrimaryKeySortedIndexScan rejects full-text search (src/paimon/core/table/source/primary_key_sorted_index_scan.cpp:288-290).
  • There is no scan, split, bucket search or read for full-text payloads.
  • Indexed scores are not propagated through the primary-key physical-position read path (the TODO in src/paimon/core/operation/raw_file_split_read.cpp:85-89).

Java design (apache/paimon#8649, apache/paimon#8652, apache/paimon#8659, apache/paimon#8844, apache/paimon#9060, apache/paimon#9184):

  • Dispatch. FullTextSearchBuilderImpl routes to the primary-key path when the table is not a data-evolution table and pk-full-text.index.columns covers the column. A non-partition filter on that path is rejected with Primary-key full-text search does not support non-partition filters yet.
  • PrimaryKeyFullTextScan:
    1. Plan the primary-key batch scan with the partition filter, pinned to the snapshot.
    2. Scan the index manifest for full-text payloads that have source metadata and the definition's field id. All entries must be ADD.
    3. Group data splits by (partition, bucket), skipping bucket < 0.
    4. Keep eligible files with their aligned DeletionFiles, and resolve current payloads with PkFullTextBucketIndexState#fromActiveDataFiles.
    5. Emit one split per bucket.
  • PrimaryKeyFullTextSearchSplit holds the data split, the payload files and the uncovered data file names. Every eligible file is either covered by exactly one payload or listed as uncovered.
  • PrimaryKeyFullTextBucketSearch (searchRankingsAsync):
    • For each payload, lay out its source files in order to get their row offsets.
    • If a source is inactive or has deletions, the include set is the active ranges minus deleted positions. A payload whose include set is empty is skipped.
    • Call visitFullTextSearch(new FullTextSearch(column, query, limit).withIncludeRowIds(include)).
    • Map hits to PrimaryKeySearchPosition(partition, bucket, fileName, rowId - offset, score), sorted by score descending, then file name, then position.
  • PrimaryKeyFullTextRead:
    • Requires limit > 0, and supports only full-text-index.search-mode = fast; full/detail throw UnsupportedOperationException.
    • Searches each split asynchronously on the global-index.thread-num executor, submitting from the caller thread.
    • Takes the global top-k with PrimaryKeySearchRanker#topKByScore.
    • Returns PrimaryKeyScoredResult. It turns into one indexed split per data file, with row ranges, scores and that file's DeletionFile, which are read by position.
    • Uncovered data files are not searched in fast mode.
Solution
  • Port PrimaryKeyFullTextScan, PrimaryKeyFullTextSearchSplit (serializable), PrimaryKeyFullTextBucketSearch and PrimaryKeyFullTextRead.
  • Port the shared primitives PrimaryKeySearchPosition, PrimaryKeySearchRanker#topKByScore and PrimaryKeyScoredResult. Hybrid search (#407) will reuse them.
  • Propagate _INDEX_SCORE through the primary-key positional read path.
  • Add primary-key dispatch to the table-level full-text builder from #404.
  • Add tests aligned with Java PrimaryKeyFullTextScanTest, PrimaryKeyFullTextReadTest, PrimaryKeyFullTextSearchTest, PrimaryKeyFullTextBucketSearchTest and NativePrimaryKeyFullTextIndexTest:
    • deletion vectors
    • several buckets and levels
    • uncovered files
    • global top-k
    • partition filters
    • rejection of non-partition filters and non-fast modes
    • archives written by Java
Anything else?

Depends on #404 and #409.

Are you willing to submit a PR?
  • I'm willing to submit a PR!
Lingua principale
C++
Stelle
65
Fork
31
Merge medio
1g 23h
PR unite (30g)
64

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/paimon-cpp

Tutte le issue di apache/paimon-cpp

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.