[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer
メンテナーはふだん 1 日以内に返信
評価
この issue はまだ評価されていません。
説明
Search before asking
- I searched in the issues and found nothing similar.
Motivation
After a Parquet data file is closed, Paimon C++ gets the file's column statistics by opening the file it has just written and reading its footer back:
DataFileWriter::GetResult()→GetFieldStats()→stats_extractor_->Extract(fs_, path_, pool_)(src/paimon/core/io/data_file_writer.cpp:57,:82).KeyValueDataFileWriter::GetFieldStats()does the same for primary-key tables (src/paimon/core/io/key_value_data_file_writer.cpp:200).ParquetStatsExtractor::ExtractWithFileInfo()(src/paimon/format/parquet/parquet_stats_extractor.cpp:257) callsFileSystem::Open(path)andInputStream::Length(), then opens aparquet::arrow::FileReader. That reader reads the file tail (Arrow 17 reads a fixed 64 KiB tail, plus a second read when the footer is larger) and Thrift-decodes the wholeFileMetaDataagain.
The writer already holds exactly this metadata. After parquet::arrow::FileWriter::Close(), FileWriter::metadata() returns the FileMetaData that was serialized into the footer. Re-reading it costs, for every written file:
- An open plus a length lookup. On Jindo/OSS this includes the
getFileStatusround trip that #331 avoids on the read side when the length is known. - One or more ranged reads of the footer.
- A full Thrift decode of the footer, whose cost grows with columns × row groups.
This happens for every data file produced by writes and by compaction, on both append and primary-key tables, before the file's DataFileMeta can be built. On high-latency storage it adds round trips for every rolled file, and the overhead weighs most when many small files are produced, for example level-0 files or many partitions/buckets each writing a file.
Velox's Parquet writer uses the in-memory metadata instead: Writer::close() returns a ParquetFileMetadata wrapping arrowContext_->writer->metadata() (velox/dwio/parquet/writer/Writer.cpp), and statistics are derived from it without re-reading the file.
Solution
Derive the stats from the metadata the writer already holds, and keep the re-read path as a fallback.
- Keep the metadata. In
ParquetFormatWriter::Finish(), afterwriter_->Close(), retainwriter_->metadata()(std::shared_ptr<parquet::FileMetaData>). - Share the conversion. Factor the part of
ParquetStatsExtractor::ExtractWithFileInfo()that runs after the footer is obtained into a helper takingconst parquet::FileMetaData&. That covers the per-row-group merge throughMergeStats,ConvertStatsToColumnStats, nested-field handling andFileInfo(num_rows). The re-read path and the in-memory path then produce stats through the same code. - Use it from the data file writers.
DataFileWriterandKeyValueDataFileWritertake the stats from the format writer when it can provide them, and fall back tostats_extractor_->Extract(fs_, path_, pool_)otherwise, for example for other formats.FormatWriterandFormatStatsExtractorare exported underinclude/paimon/format/, so the plumbing should preferably stay internal, e.g. an internal interface that the Parquet writer implements and the data file writer checks for. A new virtual method on the public classes would change their ABI; if that route is preferred, the impact should be called out. - Out of scope. Callers that extract stats from files they did not write keep re-reading the footer, e.g. migration (
src/paimon/core/migrate/file_meta_utils.cpp:83). ORC, Avro and Lance can adopt the same hook later.
Anything else?
The in-memory FileMetaData is not built the same way as a decoded footer, so equivalence has to be tested rather than assumed. In Arrow 17, FileMetaDataBuilder::Finish() (cpp/src/parquet/metadata.cc) default-constructs the FileMetaData and only calls InitSchema() and InitKeyValueMetadata(). As a result:
writer_version_is not parsed fromcreated_byand stays a defaultApplicationVersionwith an empty application name.InitColumnOrders()is not called.
Reading ColumnChunkMetaData::is_stats_set() and ApplicationVersion::HasCorrectStatistics(), both paths should accept the statistics of every column with a known sort order. An empty application name is neither parquet-cpp, parquet-mr nor unknown, and the PARQUET-251 check compares application names before versions. A test should still pin this down: write files that cover
- every supported primitive type, including DECIMAL stored as INT32/INT64, TIMESTAMP including INT96, STRING/BINARY, and FLOAT/DOUBLE with NaN;
- all-null columns;
- nested types;
- multiple row groups;
and assert that the in-memory stats equal the stats obtained by re-reading the footer.
Suggested validation:
- Unit tests for the equivalence above.
- A test with a counting
FileSystemasserting that closing a Parquet data file no longer opens it for reading. - Existing write and compaction integration tests pass unchanged. The
SimpleStatsstored inDataFileMetamust stay identical.
No storage format or protocol change.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- 主要言語
- C++
- スター
- 65
- フォーク
- 31
- 平均マージ
- 1日 14時間
- マージ済み PR(30日)
- 60
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
apache/paimon-cpp のほかの issue
-
[Feature] Warm up next data file in ConcatBatchReader対応中かも @SteNicholas が 2 日前に担当しました。 オープンenhancement
apache/paimon-cpp#419 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
[Feature] Support writing MAP<K, BLOB> fields対応中かも @SteNicholas が 4 日前に担当しました。 オープンenhancement
apache/paimon-cpp#415 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
apache/paimon-cpp#410 ·
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
apache/paimon-cpp#409 ·
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
apache/paimon-cpp#408 ·
メンテナーはふだん 1 日以内に返信
apache/paimon-cpp の issue をすべて見る
似ている issue
-
feature request
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
gavinlouuu-kpt/mib-studio-qt#517 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 73/100
EchoTools/nevr-runtime#116 ·
メンテナーはふだん 1 日以内に返信