[Feature] Resolve primary-key full-text index options and validate schemas
維護者通常 1 天內回覆
還沒有人認領這個 Issue。
評估
研究方向
Start with src/paimon/core/index/pk/primary_key_index_definitions.cpp:208-211 and src/paimon/core/schema/schema_validation.cpp:565-570, then read docs/source/user_guide/primary_key_global_index.rst. Use the listed Java validation and option-resolution tests as behavioral references. Done means full-text options are preserved, all listed validation and cross-family checks run independently, documentation is corrected, and corresponding tests cover the specified cases.
由索引模型根據 Issue 內容生成。
描述
Search before asking
- I searched in the issues and found nothing similar.
Motivation
Sub-issue of #399 (step 5: primary-key full-text index, options and validation).
Paimon C++ recognizes pk-full-text.index.columns, but two parts are missing:
- Options are dropped.
PrimaryKeyIndexDefinitions::Createbuilds theFULL_TEXTdefinition with an empty options map (src/paimon/core/index/pk/primary_key_index_definitions.cpp:208-211). Tokenizer options such asfull-text.tokenizer=jiebanever reach the index. - Validation is skipped.
SchemaValidation::ValidatePrimaryKeyBTreeIndexesreturns early whenpk-btree.index.columnsis empty (src/paimon/core/schema/schema_validation.cpp:565-570). A table that sets onlypk-full-text.index.columnstherefore gets none of the checks: column existence, column type, cross-family ownership, deletion vectors, bucket mode.docs/source/user_guide/primary_key_global_index.rstsays vector and full-text definitions "are recognized for validation", which is only true when BTree columns are also set.
Java behavior (apache/paimon#8651, apache/paimon#8672, apache/paimon#8922):
Option merging (CoreOptions#primaryKeyFullTextIndexOptions(column)):
- Start from the table options whose key starts with
full-text.. - Parse
fields.<column>.pk-full-text.index.optionsas a JSON object of string values. Reject:- anything that is not a JSON object:
<key> must be a JSON object of option key-value pairs. - an empty key:
<key> contains an empty option key. - a null value:
<key> value for key <k> must not be null.
- anything that is not a JSON object:
- Qualify keys that lack the
full-text.prefix. - Reject a key already set at table level with a different value:
<key> defines conflicting values for full-text.<k>.An equal value is accepted.
For example, {"full-text.tokenizer":"jieba","ngram.min-gram":"2"} resolves to full-text.tokenizer=jieba and full-text.ngram.min-gram=2. The full-text indexer later strips the prefix (see #400).
Schema validation (SchemaValidation#validatePrimaryKeyFullTextIndex, run when the key is present):
- Exactly one column:
pk-full-text.index.columns must contain exactly one column in the first release, but is [...]. - The column is not blank.
- The table is a primary-key table.
deletion-vectors.enabled = true, unless the merge engine isfirst-row.deletion-vectors.merge-on-read = falsewhen deletion vectors are enabled.- Fixed or postpone bucket mode:
bucket > 0orbucket = -2. - No
pk-clustering-override. - The column exists.
- The column type is
CHAR/VARCHAR/STRING. - The resolved options are valid, as in the merging rules above.
Cross-family checks apply to every primary-key index family:
- no duplicate column within one family key;
- a column can own at most one primary-key index across all families.
The maintainer factory also allows only one FULL_TEXT definition: Only one primary-key full-text index is supported.
Solution
- Resolve the full-text options for each
FULL_TEXTdefinition fromfields.<column>.pk-full-text.index.optionsand the table-levelfull-text.*options, using the rules above. - Validate primary-key index columns independently of whether BTree columns are set. Add the full-text rules and the cross-family checks.
- Fix the statement in
primary_key_global_index.rst. - Add tests aligned with Java
PrimaryKeyFullTextIndexValidationTestandPrimaryKeyIndexDefinitionsTest#testResolvesFullTextIndexOptions:- more than one column, a duplicate column, a blank column, an unknown column, an
INTcolumn - deletion vectors off, and
first-rowwith deletion vectors off - merge-on-read on
bucket = -1rejected,bucket = -2acceptedpk-clustering-override- the column already used by
pk-btree - malformed JSON
- a conflicting
tokenizer - an append table
- more than one column, a duplicate column, a blank column, an unknown column, an
Anything else?
- Java also rejects renaming, dropping, or changing the type of a primary-key index column during schema evolution (
SchemaManagerUtils#assertNotUpdatingPrimaryKeyIndexColumn). Paimon C++ has no schema-evolution API yet, so this can be added together with one. - This issue unblocks #409 and #410.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- 主要語言
- C++
- 星號
- 65
- 分支
- 31
- 平均合併
- 1 天 14 小時
- 30 天內合併 PR
- 60
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
apache/paimon-cpp 的其他 Issue
-
[Feature] Warm up next data file in ConcatBatchReader可能已有人在做 @SteNicholas 於 2 天前認領。 未關閉enhancement
apache/paimon-cpp#419 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer可能已有人在做 @SteNicholas 於 2 天前認領。 未關閉enhancement
apache/paimon-cpp#417 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
[Feature] Support writing MAP<K, BLOB> fields可能已有人在做 @SteNicholas 於 4 天前認領。 未關閉enhancement
apache/paimon-cpp#415 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
enhancement
難度 5/5 一週以上 新手友好度 35/100
apache/paimon-cpp#410 ·
維護者通常 1 天內回覆
-
enhancement
難度 5/5 一週以上 新手友好度 25/100
apache/paimon-cpp#409 ·
維護者通常 1 天內回覆
查看 apache/paimon-cpp 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 68/100
維護者通常 1 天內回覆
-
bug
難度 2/5 1-3 小時 新手友好度 73/100
EchoTools/nevr-runtime#116 ·
維護者通常 1 天內回覆
-
code-quality libc++
難度 1/5 1 小時以內 新手友好度 82/100
llvm/llvm-project#229284 ·
維護者通常 1 天內回覆
-
test-issue
難度 2/5 1-3 小時 新手友好度 82/100
llvm/offload-test-suite#1557 ·
維護者通常 1 天內回覆
-
enhancement
難度 2/5 1-3 小時 新手友好度 72/100
維護者通常 1 天內回覆