[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer
I maintainer di solito rispondono entro 1 giorno
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Search before asking
- I searched in the issues and found nothing similar.
Motivation
After a Parquet data file is closed, Paimon C++ gets the file's column statistics by opening the file it has just written and reading its footer back:
DataFileWriter::GetResult()→GetFieldStats()→stats_extractor_->Extract(fs_, path_, pool_)(src/paimon/core/io/data_file_writer.cpp:57,:82).KeyValueDataFileWriter::GetFieldStats()does the same for primary-key tables (src/paimon/core/io/key_value_data_file_writer.cpp:200).ParquetStatsExtractor::ExtractWithFileInfo()(src/paimon/format/parquet/parquet_stats_extractor.cpp:257) callsFileSystem::Open(path)andInputStream::Length(), then opens aparquet::arrow::FileReader. That reader reads the file tail (Arrow 17 reads a fixed 64 KiB tail, plus a second read when the footer is larger) and Thrift-decodes the wholeFileMetaDataagain.
The writer already holds exactly this metadata. After parquet::arrow::FileWriter::Close(), FileWriter::metadata() returns the FileMetaData that was serialized into the footer. Re-reading it costs, for every written file:
- An open plus a length lookup. On Jindo/OSS this includes the
getFileStatusround trip that #331 avoids on the read side when the length is known. - One or more ranged reads of the footer.
- A full Thrift decode of the footer, whose cost grows with columns × row groups.
This happens for every data file produced by writes and by compaction, on both append and primary-key tables, before the file's DataFileMeta can be built. On high-latency storage it adds round trips for every rolled file, and the overhead weighs most when many small files are produced, for example level-0 files or many partitions/buckets each writing a file.
Velox's Parquet writer uses the in-memory metadata instead: Writer::close() returns a ParquetFileMetadata wrapping arrowContext_->writer->metadata() (velox/dwio/parquet/writer/Writer.cpp), and statistics are derived from it without re-reading the file.
Solution
Derive the stats from the metadata the writer already holds, and keep the re-read path as a fallback.
- Keep the metadata. In
ParquetFormatWriter::Finish(), afterwriter_->Close(), retainwriter_->metadata()(std::shared_ptr<parquet::FileMetaData>). - Share the conversion. Factor the part of
ParquetStatsExtractor::ExtractWithFileInfo()that runs after the footer is obtained into a helper takingconst parquet::FileMetaData&. That covers the per-row-group merge throughMergeStats,ConvertStatsToColumnStats, nested-field handling andFileInfo(num_rows). The re-read path and the in-memory path then produce stats through the same code. - Use it from the data file writers.
DataFileWriterandKeyValueDataFileWritertake the stats from the format writer when it can provide them, and fall back tostats_extractor_->Extract(fs_, path_, pool_)otherwise, for example for other formats.FormatWriterandFormatStatsExtractorare exported underinclude/paimon/format/, so the plumbing should preferably stay internal, e.g. an internal interface that the Parquet writer implements and the data file writer checks for. A new virtual method on the public classes would change their ABI; if that route is preferred, the impact should be called out. - Out of scope. Callers that extract stats from files they did not write keep re-reading the footer, e.g. migration (
src/paimon/core/migrate/file_meta_utils.cpp:83). ORC, Avro and Lance can adopt the same hook later.
Anything else?
The in-memory FileMetaData is not built the same way as a decoded footer, so equivalence has to be tested rather than assumed. In Arrow 17, FileMetaDataBuilder::Finish() (cpp/src/parquet/metadata.cc) default-constructs the FileMetaData and only calls InitSchema() and InitKeyValueMetadata(). As a result:
writer_version_is not parsed fromcreated_byand stays a defaultApplicationVersionwith an empty application name.InitColumnOrders()is not called.
Reading ColumnChunkMetaData::is_stats_set() and ApplicationVersion::HasCorrectStatistics(), both paths should accept the statistics of every column with a known sort order. An empty application name is neither parquet-cpp, parquet-mr nor unknown, and the PARQUET-251 check compares application names before versions. A test should still pin this down: write files that cover
- every supported primitive type, including DECIMAL stored as INT32/INT64, TIMESTAMP including INT96, STRING/BINARY, and FLOAT/DOUBLE with NaN;
- all-null columns;
- nested types;
- multiple row groups;
and assert that the in-memory stats equal the stats obtained by re-reading the footer.
Suggested validation:
- Unit tests for the equivalence above.
- A test with a counting
FileSystemasserting that closing a Parquet data file no longer opens it for reading. - Existing write and compaction integration tests pass unchanged. The
SimpleStatsstored inDataFileMetamust stay identical.
No storage format or protocol change.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- Lingua principale
- C++
- Stelle
- 65
- Fork
- 31
- Merge medio
- 1g 14h
- PR unite (30g)
- 60
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di apache/paimon-cpp
-
[Feature] Warm up next data file in ConcatBatchReaderForse già presa @SteNicholas l’ha presa 1 giorno fa. Apertaenhancement
apache/paimon-cpp#419 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
[Feature] Support writing MAP<K, BLOB> fieldsForse già presa @SteNicholas l’ha presa 3 giorni fa. Apertaenhancement
apache/paimon-cpp#415 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
apache/paimon-cpp#410 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
apache/paimon-cpp#409 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
apache/paimon-cpp#408 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di apache/paimon-cpp
Issue simili
-
code-quality libc++
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
llvm/llvm-project#229284 ·
I maintainer di solito rispondono entro 1 giorno
-
test-issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
llvm/offload-test-suite#1557 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
iOS: hidden scale bar invalidates its intrinsic content size on every layout pass of MLNMapViewAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
maplibre/maplibre-native#4723 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
HarbourMasters/Shipwright#7320 ·
I maintainer di solito rispondono entro 1 giorno