Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

[Feature] Search primary-key full-text indexes

Aberta
#410 0 comentários 0 reações 0 responsáveis Ver no GitHub

Mantenedores costumam responder em até 1 dia

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
5/5
Tempo estimado
Mais de uma semana
Facilidade para iniciantes
35/100
Tipo de issue
Funcionalidade
Clareza
Razoavelmente clara
Status de atividade
Ativa
Stack de tecnologia
cpp
Domínio
backend, databases, search

Direção de pesquisa

Start with the rejection in src/paimon/core/table/source/primary_key_sorted_index_scan.cpp and the TODO in src/paimon/core/operation/raw_file_split_read.cpp, then review dependencies #404 and #409. Use the Java-aligned PrimaryKeyFullTextScanTest, PrimaryKeyFullTextReadTest, PrimaryKeyFullTextSearchTest, PrimaryKeyFullTextBucketSearchTest, and NativePrimaryKeyFullTextIndexTest as behavioral guides. Done means primary-key full-text searches plan, rank, read, and propagate scores with the listed filter, mode, deletion, bucket, and archive cases covered.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 5: primary-key full-text index, read side).

Paimon C++ cannot search primary-key full-text indexes:

  • PrimaryKeySortedIndexScan rejects full-text search (src/paimon/core/table/source/primary_key_sorted_index_scan.cpp:288-290).
  • There is no scan, split, bucket search or read for full-text payloads.
  • Indexed scores are not propagated through the primary-key physical-position read path (the TODO in src/paimon/core/operation/raw_file_split_read.cpp:85-89).

Java design (apache/paimon#8649, apache/paimon#8652, apache/paimon#8659, apache/paimon#8844, apache/paimon#9060, apache/paimon#9184):

  • Dispatch. FullTextSearchBuilderImpl routes to the primary-key path when the table is not a data-evolution table and pk-full-text.index.columns covers the column. A non-partition filter on that path is rejected with Primary-key full-text search does not support non-partition filters yet.
  • PrimaryKeyFullTextScan:
    1. Plan the primary-key batch scan with the partition filter, pinned to the snapshot.
    2. Scan the index manifest for full-text payloads that have source metadata and the definition's field id. All entries must be ADD.
    3. Group data splits by (partition, bucket), skipping bucket < 0.
    4. Keep eligible files with their aligned DeletionFiles, and resolve current payloads with PkFullTextBucketIndexState#fromActiveDataFiles.
    5. Emit one split per bucket.
  • PrimaryKeyFullTextSearchSplit holds the data split, the payload files and the uncovered data file names. Every eligible file is either covered by exactly one payload or listed as uncovered.
  • PrimaryKeyFullTextBucketSearch (searchRankingsAsync):
    • For each payload, lay out its source files in order to get their row offsets.
    • If a source is inactive or has deletions, the include set is the active ranges minus deleted positions. A payload whose include set is empty is skipped.
    • Call visitFullTextSearch(new FullTextSearch(column, query, limit).withIncludeRowIds(include)).
    • Map hits to PrimaryKeySearchPosition(partition, bucket, fileName, rowId - offset, score), sorted by score descending, then file name, then position.
  • PrimaryKeyFullTextRead:
    • Requires limit > 0, and supports only full-text-index.search-mode = fast; full/detail throw UnsupportedOperationException.
    • Searches each split asynchronously on the global-index.thread-num executor, submitting from the caller thread.
    • Takes the global top-k with PrimaryKeySearchRanker#topKByScore.
    • Returns PrimaryKeyScoredResult. It turns into one indexed split per data file, with row ranges, scores and that file's DeletionFile, which are read by position.
    • Uncovered data files are not searched in fast mode.
Solution
  • Port PrimaryKeyFullTextScan, PrimaryKeyFullTextSearchSplit (serializable), PrimaryKeyFullTextBucketSearch and PrimaryKeyFullTextRead.
  • Port the shared primitives PrimaryKeySearchPosition, PrimaryKeySearchRanker#topKByScore and PrimaryKeyScoredResult. Hybrid search (#407) will reuse them.
  • Propagate _INDEX_SCORE through the primary-key positional read path.
  • Add primary-key dispatch to the table-level full-text builder from #404.
  • Add tests aligned with Java PrimaryKeyFullTextScanTest, PrimaryKeyFullTextReadTest, PrimaryKeyFullTextSearchTest, PrimaryKeyFullTextBucketSearchTest and NativePrimaryKeyFullTextIndexTest:
    • deletion vectors
    • several buckets and levels
    • uncovered files
    • global top-k
    • partition filters
    • rejection of non-partition filters and non-fast modes
    • archives written by Java
Anything else?

Depends on #404 and #409.

Are you willing to submit a PR?
  • I'm willing to submit a PR!
Linguagem predominante
C++
Estrelas
65
Forks
31
Merge médio
1d 14h
PRs com merge (30d)
60

Preparar o ambiente

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de apache/paimon-cpp

Todas as issues de apache/paimon-cpp

Issues semelhantes

Mais issues de C++

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.