[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer
Maintainer antworten meist innerhalb von 1 Tag
Bewertung
Dieses Issue wurde noch nicht bewertet.
Beschreibung
Search before asking
- I searched in the issues and found nothing similar.
Motivation
After a Parquet data file is closed, Paimon C++ gets the file's column statistics by opening the file it has just written and reading its footer back:
DataFileWriter::GetResult()→GetFieldStats()→stats_extractor_->Extract(fs_, path_, pool_)(src/paimon/core/io/data_file_writer.cpp:57,:82).KeyValueDataFileWriter::GetFieldStats()does the same for primary-key tables (src/paimon/core/io/key_value_data_file_writer.cpp:200).ParquetStatsExtractor::ExtractWithFileInfo()(src/paimon/format/parquet/parquet_stats_extractor.cpp:257) callsFileSystem::Open(path)andInputStream::Length(), then opens aparquet::arrow::FileReader. That reader reads the file tail (Arrow 17 reads a fixed 64 KiB tail, plus a second read when the footer is larger) and Thrift-decodes the wholeFileMetaDataagain.
The writer already holds exactly this metadata. After parquet::arrow::FileWriter::Close(), FileWriter::metadata() returns the FileMetaData that was serialized into the footer. Re-reading it costs, for every written file:
- An open plus a length lookup. On Jindo/OSS this includes the
getFileStatusround trip that #331 avoids on the read side when the length is known. - One or more ranged reads of the footer.
- A full Thrift decode of the footer, whose cost grows with columns × row groups.
This happens for every data file produced by writes and by compaction, on both append and primary-key tables, before the file's DataFileMeta can be built. On high-latency storage it adds round trips for every rolled file, and the overhead weighs most when many small files are produced, for example level-0 files or many partitions/buckets each writing a file.
Velox's Parquet writer uses the in-memory metadata instead: Writer::close() returns a ParquetFileMetadata wrapping arrowContext_->writer->metadata() (velox/dwio/parquet/writer/Writer.cpp), and statistics are derived from it without re-reading the file.
Solution
Derive the stats from the metadata the writer already holds, and keep the re-read path as a fallback.
- Keep the metadata. In
ParquetFormatWriter::Finish(), afterwriter_->Close(), retainwriter_->metadata()(std::shared_ptr<parquet::FileMetaData>). - Share the conversion. Factor the part of
ParquetStatsExtractor::ExtractWithFileInfo()that runs after the footer is obtained into a helper takingconst parquet::FileMetaData&. That covers the per-row-group merge throughMergeStats,ConvertStatsToColumnStats, nested-field handling andFileInfo(num_rows). The re-read path and the in-memory path then produce stats through the same code. - Use it from the data file writers.
DataFileWriterandKeyValueDataFileWritertake the stats from the format writer when it can provide them, and fall back tostats_extractor_->Extract(fs_, path_, pool_)otherwise, for example for other formats.FormatWriterandFormatStatsExtractorare exported underinclude/paimon/format/, so the plumbing should preferably stay internal, e.g. an internal interface that the Parquet writer implements and the data file writer checks for. A new virtual method on the public classes would change their ABI; if that route is preferred, the impact should be called out. - Out of scope. Callers that extract stats from files they did not write keep re-reading the footer, e.g. migration (
src/paimon/core/migrate/file_meta_utils.cpp:83). ORC, Avro and Lance can adopt the same hook later.
Anything else?
The in-memory FileMetaData is not built the same way as a decoded footer, so equivalence has to be tested rather than assumed. In Arrow 17, FileMetaDataBuilder::Finish() (cpp/src/parquet/metadata.cc) default-constructs the FileMetaData and only calls InitSchema() and InitKeyValueMetadata(). As a result:
writer_version_is not parsed fromcreated_byand stays a defaultApplicationVersionwith an empty application name.InitColumnOrders()is not called.
Reading ColumnChunkMetaData::is_stats_set() and ApplicationVersion::HasCorrectStatistics(), both paths should accept the statistics of every column with a known sort order. An empty application name is neither parquet-cpp, parquet-mr nor unknown, and the PARQUET-251 check compares application names before versions. A test should still pin this down: write files that cover
- every supported primitive type, including DECIMAL stored as INT32/INT64, TIMESTAMP including INT96, STRING/BINARY, and FLOAT/DOUBLE with NaN;
- all-null columns;
- nested types;
- multiple row groups;
and assert that the in-memory stats equal the stats obtained by re-reading the footer.
Suggested validation:
- Unit tests for the equivalence above.
- A test with a counting
FileSystemasserting that closing a Parquet data file no longer opens it for reading. - Existing write and compaction integration tests pass unchanged. The
SimpleStatsstored inDataFileMetamust stay identical.
No storage format or protocol change.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- Vorherrschende Sprache
- C++
- Sterne
- 65
- Forks
- 31
- Ø Merge
- 1 T. 14 Std.
- Gemergte PRs (30 T.)
- 60
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus apache/paimon-cpp
-
[Feature] Warm up next data file in ConcatBatchReaderEvtl. vergeben @SteNicholas hat das vor 2 Tagen übernommen. Offenenhancement
apache/paimon-cpp#419 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 1 Tag
-
[Feature] Support writing MAP<K, BLOB> fieldsEvtl. vergeben @SteNicholas hat das vor 4 Tagen übernommen. Offenenhancement
apache/paimon-cpp#415 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 1 Tag
-
enhancement
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
apache/paimon-cpp#410 ·
Maintainer antworten meist innerhalb von 1 Tag
-
enhancement
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 25/100
apache/paimon-cpp#409 ·
Maintainer antworten meist innerhalb von 1 Tag
-
enhancement
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
apache/paimon-cpp#408 ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in apache/paimon-cpp
Ähnliche Issues
-
area:runtime good first issue
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
WATonomous/wato_f1tenth#39 ·
-
[APP BUG]: Sorting by name after searching can bring up irrelevant resultsEvtl. vergeben Ein verknüpfter Pull Request ist offen oder bereits gemergt. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
shadps4-emu/shadps4-qtlauncher#465 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 85/100
duckdb/duckdb-excel#104 ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 85/100
lxqt/lxqt-powermanagement#495 ·
-
bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
Qiskit/qiskit-aer#2466 ·