Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Feature] Maintain primary-key full-text index archives on write and compaction

Abierto
#409 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
25/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
cpp
Área
databases

Línea de trabajo

Start with src/paimon/core/index/pk/bucketed_primary_key_index_maintainer.cpp and the level-planning code in src/paimon/core/index/pksorted/, after reviewing dependencies #400 and #408. Use the Java-aligned PkFullTextIndexFileTest, PkFullTextDataFileReaderTest, PkFullTextBucketIndexStateTest, and BucketedFullTextIndexMaintainerTest as behavioral references; done means archives are maintained through restore, commit, compaction, and abort, with the listed cross-language tests passing.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Sub-issue of #399 (step 5: primary-key full-text index, write side).

BucketedPrimaryKeyIndexMaintainer keeps only BTree definitions (src/paimon/core/index/pk/bucketed_primary_key_index_maintainer.cpp:128-132), and restore keeps only BTree payloads (IsPrimaryKeyBTreePayload). This has two effects:

  • Paimon C++ writers never build primary-key full-text archives.
  • When C++ compacts a table that already has Java-written archives, those archives are neither rebuilt nor retired. They keep pointing at data files that compaction has removed, and the new compacted files are not indexed.

Java model (apache/paimon#8649, apache/paimon#8651, apache/paimon#8672, apache/paimon#8992):

  • One archive per level. Each (partition, bucket, non-zero data level) has exactly one immutable archive. It covers all eligible files of that level, sorted by file name.
  • Eligible files. A file is eligible when fileSource == COMPACT && level > 0 (PrimaryKeyIndexSourcePolicy).
  • Row ids. An archive's row ids are the concatenated physical row positions of its source files, including rows deleted by deletion vectors and null rows.
  • PkFullTextIndexFile builds one archive:
    • index type full-text;
    • every source has the same level > 0 and a positive row count;
    • it calls GlobalIndexSingleColumnWriter#write(text, sourceOffset + rowPos) for each row;
    • finish() must return exactly one entry whose row count equals the total source row count.
    • The resulting IndexFileMeta has GlobalIndexMeta(0, total - 1, fieldId, null, indexMeta, sourceMeta):
      • indexMeta is the flat JSON of the prefix-stripped options;
      • sourceMeta is PrimaryKeyIndexSourceMeta(level, sourceFiles).
    • The file is named index-{uuid}-{N} under the index directory, or in the bucket directory when index-file-in-data-file-dir is set.
  • PkFullTextDataFileReader reads one text value per physical row, with no deletion-vector filtering. PkFullTextIndexBuilder builds an archive for one file or a list of files.
  • PkFullTextBucketIndexState#fromActiveDataFiles classifies payloads:
    • A payload is current only if its (level, ordered source files) exactly matches the level's eligible active files and its row counts match.
    • A payload that fails this is stale. So is every payload of a level that has more than one match, and any payload whose metadata cannot be parsed.
  • Level planning (PrimaryKeyIndexLevels, shared with the sorted indexes) picks the lowest level whose payload is missing or out of date. A plan with no source files removes the payload.
  • BucketedFullTextIndexMaintainer:
    • Restore: stale payloads are retired and emitted as deletions in the next commit.
    • prepareCommit(append, compact, waitCompaction):
      1. Applies the data transition: removes compactBefore files and adds eligible compactAfter files.
      2. Finishes or starts a single background build.
      3. Atomically replaces the level's archive when the build is still valid (canAccept); otherwise deletes the generated file.
      4. Routes the result to the compact increment if there was a compact transition, and to the append increment otherwise.
      5. Rolls back and deletes generated files on failure. The returned commit carries an abort hook.
    • A failed build is not retried within the same call. The next prepareCommit plans again.
  • Wiring:
    • BucketedPrimaryKeyIndexMaintainer.Factory creates the full-text maintainer for fixed-bucket writes and for postpone-bucket compaction (apache/paimon#8992).
    • IndexFileHandler#pkFullTextIndex(partition, bucket) provides the index file.
    • Writers are not closed while a build is pending.
Solution
  • Port the classes above.
  • Reuse the existing C++ PrimaryKeyIndexSourceMeta, PrimaryKeyIndexSourcePolicy and PrimaryKeyIndexSourceFile. Extract the level planning currently embedded in src/paimon/core/index/pksorted/ so it can be shared.
  • Wire the full-text maintainer into BucketedPrimaryKeyIndexMaintainer: restore, prepareCommit, abort, and merging the increments. Create it through the full-text indexer from #400, using the options resolved in #408.
  • Add tests aligned with Java PkFullTextIndexFileTest, PkFullTextDataFileReaderTest, PkFullTextBucketIndexStateTest and BucketedFullTextIndexMaintainerTest, plus:
    • C++ compaction of a table with Java-written primary-key full-text archives; stale archives are retired and new ones built;
    • Java reading archives written by C++.
Anything else?
  • Depends on #400 and #408.
  • The realtime path still rejects primary-key global indexes (src/paimon/core/utils/primary_key_table_utils.cpp:136-141). That is out of scope here.
Are you willing to submit a PR?
  • I'm willing to submit a PR!
Lenguaje dominante
C++
Estrellas
66
Forks
33
Merge medio
1 d 14 h
PR fusionados (30 d)
57

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de apache/paimon-cpp

Todos los issues de apache/paimon-cpp

Issues similares

Más issues de C++

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.