Add batch/columnar write API for Arrow VectorSchemaRoot (Java parity with C++/Python)
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 30/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
- 技术栈
- java
调研方向
首先阅读 parquet-arrow 模块以及名为 PageWriteStore、PageWriter、BytesInput 和 ColumnChunkPageWriteStore 的 API。拟议的工作分为五个实现阶段,首先为非 nullable 的固定宽度 PLAIN 列实现零拷贝写入;完成情况应根据所选阶段进行评判,完整范围包括对 nullable、可变宽度和字典的支持。
由索引模型根据 Issue 内容生成。
描述
Problem
parquet-java has no public API to write Arrow VectorSchemaRoot to Parquet files. Callers must construct row objects and feed them through ParquetWriter<T>.write(T) one at a time. Arrow C++ and PyArrow support this natively via parquet::WriteTable() / pq.write_table().
Related prior discussion:
- #2264 (open since 2019)
- #3353
Motivation
Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy ParquetWriter's input API:
- Apache Iceberg (
FileAppender<Record>) — iceberg#17748 - Apache Fluss (Arrow-native streaming storage) — fluss#4047
- Apache Paimon (worked around this by building their own
paimon-arrowwriter that bypassesparquet-java's row API entirely)
Key Insight
For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a BytesInput and passing it to PageWriter.writePage().
For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).
Proposed Design
A new ArrowParquetWriter in the parquet-arrow module that writes pages directly to PageWriter rather than going through RecordConsumer/ColumnWriter:
ArrowParquetWriter.writeBatch(VectorSchemaRoot)
→ per column: selects optimal write strategy
→ writes pages to PageWriter via getPageWriter(ColumnDescriptor)
→ manages row groups via ParquetFileWriter
Per-column strategy selection (best to worst):
- Zero-copy (non-null + fixed-width + PLAIN): wrap Arrow buffer as
BytesInputdirectly. RL/DL as single-value RLE runs. Stats from sequential buffer scan. - Bulk-copy with nulls (nullable + fixed-width + PLAIN): scan validity bitmap for non-null runs, bulk-copy each, emit DL as RLE runs per null-transition.
- Variable-width rewrite (PLAIN + string/binary): single-pass transformation of Arrow offset+data buffers to Parquet's length-prefixed format.
- Dictionary: map Arrow's
DictionaryEncodedVectorto Parquet dictionary pages, or build dictionary from plain vector. - Fallback: per-value through
ValuesWriter(same cost as today, for unsupported encoding/type combos).
All verified public APIs exist for this:
PageWriteStore.getPageWriter(ColumnDescriptor)— get column page writer directlyPageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings)— write pre-encoded pagesBytesInput.from(ByteBuffer)— zero-copy buffer wrappingRunLengthBitPackingHybridEncoder— encode RL/DL levelsColumnChunkPageWriteStore.flushToFileWriter()— flush pages to file
Why Not Extend ParquetWriter?
ParquetWriter.write(T) calls InternalParquetRecordWriter.write() which increments recordCount by 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manage ParquetFileWriter and row groups directly.
Scope
- Parquet file format unchanged
- Existing
ParquetWriter<T>API unchanged parquet-arrowgainsparquet-hadoopas a compile dependency (forParquetFileWriter,ColumnChunkPageWriteStore)- No Hadoop runtime dependency for uncompressed output (compression requires caller-supplied
CompressionCodecFactory) - Flat schemas (primitive columns) first; nested types deferred
Implementation Phases
- Core infrastructure + zero-copy for non-null fixed-width PLAIN columns
- Nullable column support (validity bitmap scanning + run-based bulk copy)
- Variable-width types (string/binary offset rewriting)
- Dictionary encoding support
- Bloom filter, nested types, advanced features
- 主要语言
- Java
- 星标
- 3.1k
- 派生
- 1.6k
- 平均合并
- 4 天 12 小时
- 30 天内合并 PR
- 28
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/parquet-java 的其他 Issue
-
Type: bug
难度 2/5 1-3 小时 新手友好度 68/100
apache/parquet-java#3792 ·
-
难度 2/5 1-3 小时 新手友好度 82/100
apache/parquet-java#3767 ·
-
难度 2/5 1-3 小时 新手友好度 72/100
apache/parquet-java#3695 · 1 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
apache/parquet-java#3667 ·
-
Type: bug
难度 2/5 1-3 小时 新手友好度 76/100
apache/parquet-java#3587 ·
查看 apache/parquet-java 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 65/100
-
bug
难度 2/5 1-3 小时 新手友好度 75/100
-
难度 2/5 1-3 小时 新手友好度 75/100
elastic/gradle-plugins#157 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
cryptomator/hub#497 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
johanhaleby/occurrent#1120 ·