Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

[Feature] Support shared-shredding storage layout for MAP columns

Geschlossen
#220 0 Kommentare 0 Reaktionen 1 zugewiesene Person Auf GitHub ansehen

Maintainer antworten meist innerhalb von 1 Tag

@lszskye arbeitet bereits daran.

Seit 21.9.2026.

Bewertung

Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Anfängerfreundlichkeit
35/100
Issue-Typ
Feature
Klarheit
Größtenteils klar
Aktivitätsstatus
Aktiv
Tech-Stack
cpp
Bereich
databases

Rechercherichtung

Beginne mit PIP-43 und verfolge die bestehenden Schnittstellen FormatWriter, FileBatchReader, AppendOnlyWriter, LeafPredicate und PredicateConverter. Bilde die Schreib- und Lesepfade ab, bevor du die Komponenten für Schema-Konvertierung, Metadaten, Allokation, Prädikatsübersetzung und Rekonstruktion implementierst. Fertig bedeutet, dass shared-shredding-MAP-Spalten beim Schreiben, Lesen, Filtern, bei Overflow und bei unterschiedlichen K-Werten über Dateien hinweg funktionieren.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

In time-series / IoT / observability workloads, a common pattern is storing free-schema fields in a MAP<STRING, T> column (e.g. metrics MAP<STRING, DOUBLE>). The default MAP storage (two KV arrays) provides:

  • No per-key columnar access
  • No per-key statistics
  • No predicate pushdown on individual keys

This makes queries like SELECT ext_map['usage'] FROM metrics WHERE ext_map['usage'] > 30 scan the entire MAP column — extremely inefficient when only 1–3 keys out of thousands are needed per query.

The PIP-43: Columnar Storage Optimization for MAP Type in Paimon proposes a new shared-shredding storage layout that stores MAP values in K reusable physical columns within a Struct, achieving near-full columnar access with per-key statistics and predicate pushdown — without changing the logical type (MAP<STRING, T>).

Solution
Physical Layout

Each MAP<STRING, T> column configured with fields.<column>.map.storage-layout = shared-shredding is physically stored as:

STRUCT<
  __field_mapping: FixedSizeList<Int32, K>,   -- per-row: which field_id each col holds
  __col_0: T, __col_1: T, ..., __col_{K-1}: T,  -- reusable typed columns
  __overflow: MAP<INT32, T>                    -- rare fallback for rows with > K fields
>

fields.<column>.map.shared-shredding.max-columns controls K_max, and fields.<column>.map.shared-shredding.column-placement-policy controls column placement.

File metadata (footer) stores: field name↔id dictionary, field_id→physical column set S, overflow set O, K, and max row width.

Write Path
  1. Schema conversion utilities — Logical MAP → physical Struct schema rewriting; metadata serialization/deserialization; shared-shredding column detection via field metadata marker.

  2. FormatWriter::AddMetadata — New virtual method (default no-op) for writing key-value metadata to file footer before Finish(). Parquet implementation calls AddKeyValueMetadata.

  3. Column allocator — Per-row slot allocator that maps field IDs to up to K physical columns and sends the rest to overflow. Placement policy is configurable (plain, sequential, lru; default plain). Accumulates file-level statistics (S, O, max row width).

  4. Logical→physical batch converter — Parses logical MAP, encodes field names to integer IDs (file-level dictionary), invokes allocator per row, assembles physical Struct array.

  5. Writer integration — Extended DataFileWriter that performs conversion before writing + injects metadata on close. AppendOnlyWriter detects shared-shredding columns and routes accordingly. Cross-file K adaptation (P99 of recent max row widths, capped by K_max).

Read Path
  1. File metadata parsing — Parse shared-shredding metadata from file footer (dictionary, S, O, K). New GetFileKeyValueMetadata() method on FileBatchReader with Parquet implementation.

  2. Predicate translation — Translate logical predicates on MAP keys into conservative OR predicates over physical sub-columns. Requires extending LeafPredicate to support nested field paths and updating PredicateConverter to emit nested FieldRef.

  3. Read planning — At SetReadSchema time: look up which physical columns to read (from S), decide whether __overflow is needed (from O), translate predicates, and pass the physical schema + physical predicate down to the inner FileBatchReader unchanged.

  4. Batch reconstruction — After NextBatch: read __field_mapping per row to identify which column holds which field (fine-grained filter), gather values into logical MAP<STRING, T>. Merge overflow when needed. Correctness relies on per-row __field_mapping, not on pushdown precision.

  5. Reader integration — A wrapper reader (implements FileBatchReader) sits between the upper layer and the format-level reader. Per-file instance. Compatible with varying K across files. Orthogonal to DataEvolutionFileReader (schema evolution).

Anything else?

No response

Are you willing to submit a PR?
  • I'm willing to submit a PR!
Vorherrschende Sprache
C++
Sterne
65
Forks
31
Ø Merge
1 T. 14 Std.
Gemergte PRs (30 T.)
60

Entwicklungsumgebung

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus apache/paimon-cpp

Alle Issues in apache/paimon-cpp

Ähnliche Issues

Weitere Issues zu C++

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.