Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Feature] Search primary-key full-text indexes

Đang mở
#410 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Tính năng
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
cpp
Lĩnh vực
backend, databases, search

Hướng nghiên cứu

Start with the rejection in src/paimon/core/table/source/primary_key_sorted_index_scan.cpp and the TODO in src/paimon/core/operation/raw_file_split_read.cpp, then review dependencies #404 and #409. Use the Java-aligned PrimaryKeyFullTextScanTest, PrimaryKeyFullTextReadTest, PrimaryKeyFullTextSearchTest, PrimaryKeyFullTextBucketSearchTest, and NativePrimaryKeyFullTextIndexTest as behavioral guides. Done means primary-key full-text searches plan, rank, read, and propagate scores with the listed filter, mode, deletion, bucket, and archive cases covered.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 5: primary-key full-text index, read side).

Paimon C++ cannot search primary-key full-text indexes:

  • PrimaryKeySortedIndexScan rejects full-text search (src/paimon/core/table/source/primary_key_sorted_index_scan.cpp:288-290).
  • There is no scan, split, bucket search or read for full-text payloads.
  • Indexed scores are not propagated through the primary-key physical-position read path (the TODO in src/paimon/core/operation/raw_file_split_read.cpp:85-89).

Java design (apache/paimon#8649, apache/paimon#8652, apache/paimon#8659, apache/paimon#8844, apache/paimon#9060, apache/paimon#9184):

  • Dispatch. FullTextSearchBuilderImpl routes to the primary-key path when the table is not a data-evolution table and pk-full-text.index.columns covers the column. A non-partition filter on that path is rejected with Primary-key full-text search does not support non-partition filters yet.
  • PrimaryKeyFullTextScan:
    1. Plan the primary-key batch scan with the partition filter, pinned to the snapshot.
    2. Scan the index manifest for full-text payloads that have source metadata and the definition's field id. All entries must be ADD.
    3. Group data splits by (partition, bucket), skipping bucket < 0.
    4. Keep eligible files with their aligned DeletionFiles, and resolve current payloads with PkFullTextBucketIndexState#fromActiveDataFiles.
    5. Emit one split per bucket.
  • PrimaryKeyFullTextSearchSplit holds the data split, the payload files and the uncovered data file names. Every eligible file is either covered by exactly one payload or listed as uncovered.
  • PrimaryKeyFullTextBucketSearch (searchRankingsAsync):
    • For each payload, lay out its source files in order to get their row offsets.
    • If a source is inactive or has deletions, the include set is the active ranges minus deleted positions. A payload whose include set is empty is skipped.
    • Call visitFullTextSearch(new FullTextSearch(column, query, limit).withIncludeRowIds(include)).
    • Map hits to PrimaryKeySearchPosition(partition, bucket, fileName, rowId - offset, score), sorted by score descending, then file name, then position.
  • PrimaryKeyFullTextRead:
    • Requires limit > 0, and supports only full-text-index.search-mode = fast; full/detail throw UnsupportedOperationException.
    • Searches each split asynchronously on the global-index.thread-num executor, submitting from the caller thread.
    • Takes the global top-k with PrimaryKeySearchRanker#topKByScore.
    • Returns PrimaryKeyScoredResult. It turns into one indexed split per data file, with row ranges, scores and that file's DeletionFile, which are read by position.
    • Uncovered data files are not searched in fast mode.
Solution
  • Port PrimaryKeyFullTextScan, PrimaryKeyFullTextSearchSplit (serializable), PrimaryKeyFullTextBucketSearch and PrimaryKeyFullTextRead.
  • Port the shared primitives PrimaryKeySearchPosition, PrimaryKeySearchRanker#topKByScore and PrimaryKeyScoredResult. Hybrid search (#407) will reuse them.
  • Propagate _INDEX_SCORE through the primary-key positional read path.
  • Add primary-key dispatch to the table-level full-text builder from #404.
  • Add tests aligned with Java PrimaryKeyFullTextScanTest, PrimaryKeyFullTextReadTest, PrimaryKeyFullTextSearchTest, PrimaryKeyFullTextBucketSearchTest and NativePrimaryKeyFullTextIndexTest:
    • deletion vectors
    • several buckets and levels
    • uncovered files
    • global top-k
    • partition filters
    • rejection of non-partition filters and non-fast modes
    • archives written by Java
Anything else?

Depends on #404 and #409.

Are you willing to submit a PR?
  • I'm willing to submit a PR!
Ngôn ngữ chính
C++
Star
65
Fork
31
Merge trung bình
1 ngày 23 giờ
Pull request đã merge (30 ngày)
64

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của apache/paimon-cpp

Tất cả issue của apache/paimon-cpp

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.