[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer
維護者通常 1 天內回覆
評估
這個 Issue 還沒有評估資料。
描述
Search before asking
- I searched in the issues and found nothing similar.
Motivation
After a Parquet data file is closed, Paimon C++ gets the file's column statistics by opening the file it has just written and reading its footer back:
DataFileWriter::GetResult()→GetFieldStats()→stats_extractor_->Extract(fs_, path_, pool_)(src/paimon/core/io/data_file_writer.cpp:57,:82).KeyValueDataFileWriter::GetFieldStats()does the same for primary-key tables (src/paimon/core/io/key_value_data_file_writer.cpp:200).ParquetStatsExtractor::ExtractWithFileInfo()(src/paimon/format/parquet/parquet_stats_extractor.cpp:257) callsFileSystem::Open(path)andInputStream::Length(), then opens aparquet::arrow::FileReader. That reader reads the file tail (Arrow 17 reads a fixed 64 KiB tail, plus a second read when the footer is larger) and Thrift-decodes the wholeFileMetaDataagain.
The writer already holds exactly this metadata. After parquet::arrow::FileWriter::Close(), FileWriter::metadata() returns the FileMetaData that was serialized into the footer. Re-reading it costs, for every written file:
- An open plus a length lookup. On Jindo/OSS this includes the
getFileStatusround trip that #331 avoids on the read side when the length is known. - One or more ranged reads of the footer.
- A full Thrift decode of the footer, whose cost grows with columns × row groups.
This happens for every data file produced by writes and by compaction, on both append and primary-key tables, before the file's DataFileMeta can be built. On high-latency storage it adds round trips for every rolled file, and the overhead weighs most when many small files are produced, for example level-0 files or many partitions/buckets each writing a file.
Velox's Parquet writer uses the in-memory metadata instead: Writer::close() returns a ParquetFileMetadata wrapping arrowContext_->writer->metadata() (velox/dwio/parquet/writer/Writer.cpp), and statistics are derived from it without re-reading the file.
Solution
Derive the stats from the metadata the writer already holds, and keep the re-read path as a fallback.
- Keep the metadata. In
ParquetFormatWriter::Finish(), afterwriter_->Close(), retainwriter_->metadata()(std::shared_ptr<parquet::FileMetaData>). - Share the conversion. Factor the part of
ParquetStatsExtractor::ExtractWithFileInfo()that runs after the footer is obtained into a helper takingconst parquet::FileMetaData&. That covers the per-row-group merge throughMergeStats,ConvertStatsToColumnStats, nested-field handling andFileInfo(num_rows). The re-read path and the in-memory path then produce stats through the same code. - Use it from the data file writers.
DataFileWriterandKeyValueDataFileWritertake the stats from the format writer when it can provide them, and fall back tostats_extractor_->Extract(fs_, path_, pool_)otherwise, for example for other formats.FormatWriterandFormatStatsExtractorare exported underinclude/paimon/format/, so the plumbing should preferably stay internal, e.g. an internal interface that the Parquet writer implements and the data file writer checks for. A new virtual method on the public classes would change their ABI; if that route is preferred, the impact should be called out. - Out of scope. Callers that extract stats from files they did not write keep re-reading the footer, e.g. migration (
src/paimon/core/migrate/file_meta_utils.cpp:83). ORC, Avro and Lance can adopt the same hook later.
Anything else?
The in-memory FileMetaData is not built the same way as a decoded footer, so equivalence has to be tested rather than assumed. In Arrow 17, FileMetaDataBuilder::Finish() (cpp/src/parquet/metadata.cc) default-constructs the FileMetaData and only calls InitSchema() and InitKeyValueMetadata(). As a result:
writer_version_is not parsed fromcreated_byand stays a defaultApplicationVersionwith an empty application name.InitColumnOrders()is not called.
Reading ColumnChunkMetaData::is_stats_set() and ApplicationVersion::HasCorrectStatistics(), both paths should accept the statistics of every column with a known sort order. An empty application name is neither parquet-cpp, parquet-mr nor unknown, and the PARQUET-251 check compares application names before versions. A test should still pin this down: write files that cover
- every supported primitive type, including DECIMAL stored as INT32/INT64, TIMESTAMP including INT96, STRING/BINARY, and FLOAT/DOUBLE with NaN;
- all-null columns;
- nested types;
- multiple row groups;
and assert that the in-memory stats equal the stats obtained by re-reading the footer.
Suggested validation:
- Unit tests for the equivalence above.
- A test with a counting
FileSystemasserting that closing a Parquet data file no longer opens it for reading. - Existing write and compaction integration tests pass unchanged. The
SimpleStatsstored inDataFileMetamust stay identical.
No storage format or protocol change.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- 主要語言
- C++
- 星號
- 65
- 分支
- 31
- 平均合併
- 1 天 14 小時
- 30 天內合併 PR
- 60
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
apache/paimon-cpp 的其他 Issue
-
[Feature] Warm up next data file in ConcatBatchReader可能已有人在做 @SteNicholas 於 4 天前認領。 未關閉enhancement
apache/paimon-cpp#419 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
[Feature] Support writing MAP<K, BLOB> fields可能已有人在做 @SteNicholas 於 6 天前認領。 未關閉enhancement
apache/paimon-cpp#415 · 已指派 1 人 ·
維護者通常 1 天內回覆
-
enhancement
難度 5/5 一週以上 新手友好度 35/100
apache/paimon-cpp#410 ·
維護者通常 1 天內回覆
-
enhancement
難度 5/5 一週以上 新手友好度 25/100
apache/paimon-cpp#409 ·
維護者通常 1 天內回覆
-
enhancement
難度 5/5 一週以上 新手友好度 35/100
apache/paimon-cpp#408 ·
維護者通常 1 天內回覆
查看 apache/paimon-cpp 的全部 Issue
相似的 Issue
-
難度 1/5 1 小時以內 新手友好度 78/100
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 74/100
EsotericSoftware/spine-runtimes#3186 ·
-
難度 2/5 1-3 小時 新手友好度 68/100
維護者通常 1 天內回覆
-
Round video messages start gray and blocky with libx264: encoder is configured for 1,000,000 fps未關閉
難度 2/5 1-3 小時 新手友好度 78/100
telegramdesktop/tdesktop#31422 ·
維護者通常 9 天內回覆
-
難度 2/5 1-3 小時 新手友好度 68/100
duckdb/duckdb-quack#299 ·
維護者通常 1 天內回覆