Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Feature] Support shared-shredding storage layout for MAP columns

Open
#220 0 comments 0 reactions 1 assignee View on GitHub

@lszskye is already working on this.

Since Sep 21, 2026.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp
Domain
databases

Research direction

Start with PIP-43 and trace the existing FormatWriter, FileBatchReader, AppendOnlyWriter, LeafPredicate, and PredicateConverter interfaces. Map the write and read paths before implementing the schema conversion, metadata, allocation, predicate translation, and reconstruction pieces. Done means shared-shredding MAP columns work across writing, reading, filtering, overflow, and varying K values across files.

Written by the indexing model from the issue text.

Description

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

In time-series / IoT / observability workloads, a common pattern is storing free-schema fields in a MAP<STRING, T> column (e.g. metrics MAP<STRING, DOUBLE>). The default MAP storage (two KV arrays) provides:

  • No per-key columnar access
  • No per-key statistics
  • No predicate pushdown on individual keys

This makes queries like SELECT ext_map['usage'] FROM metrics WHERE ext_map['usage'] > 30 scan the entire MAP column — extremely inefficient when only 1–3 keys out of thousands are needed per query.

The PIP-43: Columnar Storage Optimization for MAP Type in Paimon proposes a new shared-shredding storage layout that stores MAP values in K reusable physical columns within a Struct, achieving near-full columnar access with per-key statistics and predicate pushdown — without changing the logical type (MAP<STRING, T>).

Solution
Physical Layout

Each MAP<STRING, T> column configured with fields.<column>.map.storage-layout = shared-shredding is physically stored as:

STRUCT<
  __field_mapping: FixedSizeList<Int32, K>,   -- per-row: which field_id each col holds
  __col_0: T, __col_1: T, ..., __col_{K-1}: T,  -- reusable typed columns
  __overflow: MAP<INT32, T>                    -- rare fallback for rows with > K fields
>

fields.<column>.map.shared-shredding.max-columns controls K_max, and fields.<column>.map.shared-shredding.column-placement-policy controls column placement.

File metadata (footer) stores: field name↔id dictionary, field_id→physical column set S, overflow set O, K, and max row width.

Write Path
  1. Schema conversion utilities — Logical MAP → physical Struct schema rewriting; metadata serialization/deserialization; shared-shredding column detection via field metadata marker.

  2. FormatWriter::AddMetadata — New virtual method (default no-op) for writing key-value metadata to file footer before Finish(). Parquet implementation calls AddKeyValueMetadata.

  3. Column allocator — Per-row slot allocator that maps field IDs to up to K physical columns and sends the rest to overflow. Placement policy is configurable (plain, sequential, lru; default plain). Accumulates file-level statistics (S, O, max row width).

  4. Logical→physical batch converter — Parses logical MAP, encodes field names to integer IDs (file-level dictionary), invokes allocator per row, assembles physical Struct array.

  5. Writer integration — Extended DataFileWriter that performs conversion before writing + injects metadata on close. AppendOnlyWriter detects shared-shredding columns and routes accordingly. Cross-file K adaptation (P99 of recent max row widths, capped by K_max).

Read Path
  1. File metadata parsing — Parse shared-shredding metadata from file footer (dictionary, S, O, K). New GetFileKeyValueMetadata() method on FileBatchReader with Parquet implementation.

  2. Predicate translation — Translate logical predicates on MAP keys into conservative OR predicates over physical sub-columns. Requires extending LeafPredicate to support nested field paths and updating PredicateConverter to emit nested FieldRef.

  3. Read planning — At SetReadSchema time: look up which physical columns to read (from S), decide whether __overflow is needed (from O), translate predicates, and pass the physical schema + physical predicate down to the inner FileBatchReader unchanged.

  4. Batch reconstruction — After NextBatch: read __field_mapping per row to identify which column holds which field (fine-grained filter), gather values into logical MAP<STRING, T>. Merge overflow when needed. Correctness relies on per-row __field_mapping, not on pushdown precision.

  5. Reader integration — A wrapper reader (implements FileBatchReader) sits between the upper layer and the format-level reader. Per-file instance. Compatible with varying K across files. Orthogonal to DataEvolutionFileReader (schema evolution).

Anything else?

No response

Are you willing to submit a PR?
  • I'm willing to submit a PR!
Dominant language
C++
Stars
65
Forks
29
Avg merge
2d 4h
Merged PRs (30d)
78

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/paimon-cpp

All issues in apache/paimon-cpp

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.