[Feature] Derive Parquet data file stats from in-memory writer metadata instead of re-reading footer
Los mantenedores suelen responder en 1 día
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Search before asking
- I searched in the issues and found nothing similar.
Motivation
After a Parquet data file is closed, Paimon C++ gets the file's column statistics by opening the file it has just written and reading its footer back:
DataFileWriter::GetResult()→GetFieldStats()→stats_extractor_->Extract(fs_, path_, pool_)(src/paimon/core/io/data_file_writer.cpp:57,:82).KeyValueDataFileWriter::GetFieldStats()does the same for primary-key tables (src/paimon/core/io/key_value_data_file_writer.cpp:200).ParquetStatsExtractor::ExtractWithFileInfo()(src/paimon/format/parquet/parquet_stats_extractor.cpp:257) callsFileSystem::Open(path)andInputStream::Length(), then opens aparquet::arrow::FileReader. That reader reads the file tail (Arrow 17 reads a fixed 64 KiB tail, plus a second read when the footer is larger) and Thrift-decodes the wholeFileMetaDataagain.
The writer already holds exactly this metadata. After parquet::arrow::FileWriter::Close(), FileWriter::metadata() returns the FileMetaData that was serialized into the footer. Re-reading it costs, for every written file:
- An open plus a length lookup. On Jindo/OSS this includes the
getFileStatusround trip that #331 avoids on the read side when the length is known. - One or more ranged reads of the footer.
- A full Thrift decode of the footer, whose cost grows with columns × row groups.
This happens for every data file produced by writes and by compaction, on both append and primary-key tables, before the file's DataFileMeta can be built. On high-latency storage it adds round trips for every rolled file, and the overhead weighs most when many small files are produced, for example level-0 files or many partitions/buckets each writing a file.
Velox's Parquet writer uses the in-memory metadata instead: Writer::close() returns a ParquetFileMetadata wrapping arrowContext_->writer->metadata() (velox/dwio/parquet/writer/Writer.cpp), and statistics are derived from it without re-reading the file.
Solution
Derive the stats from the metadata the writer already holds, and keep the re-read path as a fallback.
- Keep the metadata. In
ParquetFormatWriter::Finish(), afterwriter_->Close(), retainwriter_->metadata()(std::shared_ptr<parquet::FileMetaData>). - Share the conversion. Factor the part of
ParquetStatsExtractor::ExtractWithFileInfo()that runs after the footer is obtained into a helper takingconst parquet::FileMetaData&. That covers the per-row-group merge throughMergeStats,ConvertStatsToColumnStats, nested-field handling andFileInfo(num_rows). The re-read path and the in-memory path then produce stats through the same code. - Use it from the data file writers.
DataFileWriterandKeyValueDataFileWritertake the stats from the format writer when it can provide them, and fall back tostats_extractor_->Extract(fs_, path_, pool_)otherwise, for example for other formats.FormatWriterandFormatStatsExtractorare exported underinclude/paimon/format/, so the plumbing should preferably stay internal, e.g. an internal interface that the Parquet writer implements and the data file writer checks for. A new virtual method on the public classes would change their ABI; if that route is preferred, the impact should be called out. - Out of scope. Callers that extract stats from files they did not write keep re-reading the footer, e.g. migration (
src/paimon/core/migrate/file_meta_utils.cpp:83). ORC, Avro and Lance can adopt the same hook later.
Anything else?
The in-memory FileMetaData is not built the same way as a decoded footer, so equivalence has to be tested rather than assumed. In Arrow 17, FileMetaDataBuilder::Finish() (cpp/src/parquet/metadata.cc) default-constructs the FileMetaData and only calls InitSchema() and InitKeyValueMetadata(). As a result:
writer_version_is not parsed fromcreated_byand stays a defaultApplicationVersionwith an empty application name.InitColumnOrders()is not called.
Reading ColumnChunkMetaData::is_stats_set() and ApplicationVersion::HasCorrectStatistics(), both paths should accept the statistics of every column with a known sort order. An empty application name is neither parquet-cpp, parquet-mr nor unknown, and the PARQUET-251 check compares application names before versions. A test should still pin this down: write files that cover
- every supported primitive type, including DECIMAL stored as INT32/INT64, TIMESTAMP including INT96, STRING/BINARY, and FLOAT/DOUBLE with NaN;
- all-null columns;
- nested types;
- multiple row groups;
and assert that the in-memory stats equal the stats obtained by re-reading the footer.
Suggested validation:
- Unit tests for the equivalence above.
- A test with a counting
FileSystemasserting that closing a Parquet data file no longer opens it for reading. - Existing write and compaction integration tests pass unchanged. The
SimpleStatsstored inDataFileMetamust stay identical.
No storage format or protocol change.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- Lenguaje dominante
- C++
- Estrellas
- 65
- Forks
- 31
- Merge medio
- 1 d 14 h
- PR fusionados (30 d)
- 60
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de apache/paimon-cpp
-
[Feature] Warm up next data file in ConcatBatchReaderPosiblemente ocupada @SteNicholas la tomó hace 2 días. Abiertoenhancement
apache/paimon-cpp#419 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
[Feature] Support writing MAP<K, BLOB> fieldsPosiblemente ocupada @SteNicholas la tomó hace 4 días. Abiertoenhancement
apache/paimon-cpp#415 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
apache/paimon-cpp#410 ·
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
apache/paimon-cpp#409 ·
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
apache/paimon-cpp#408 ·
Los mantenedores suelen responder en 1 día
Todos los issues de apache/paimon-cpp
Issues similares
-
feature request
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
gavinlouuu-kpt/mib-studio-qt#517 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Los mantenedores suelen responder en 1 día
-
ws_bridge: stripping format=evr for matchmaker connections can concatenate the path and queryAbiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 73/100
EchoTools/nevr-runtime#116 ·
Los mantenedores suelen responder en 1 día