Add batch/columnar write API for Arrow VectorSchemaRoot (Java parity with C++/Python)
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 30/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- java
- Domain
- data-engineering
Research direction
Start by reading the parquet-arrow module and the named PageWriteStore, PageWriter, BytesInput, and ColumnChunkPageWriteStore APIs. The proposed work spans five implementation phases, beginning with zero-copy writes for non-null fixed-width PLAIN columns; completion should be judged against the chosen phase, with the full scope including nullable, variable-width, and dictionary support.
Written by the indexing model from the issue text.
Description
Problem
parquet-java has no public API to write Arrow VectorSchemaRoot to Parquet files. Callers must construct row objects and feed them through ParquetWriter<T>.write(T) one at a time. Arrow C++ and PyArrow support this natively via parquet::WriteTable() / pq.write_table().
Related prior discussion:
- #2264 (open since 2019)
- #3353
Motivation
Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy ParquetWriter's input API:
- Apache Iceberg (
FileAppender<Record>) — iceberg#17748 - Apache Fluss (Arrow-native streaming storage) — fluss#4047
- Apache Paimon (worked around this by building their own
paimon-arrowwriter that bypassesparquet-java's row API entirely)
Key Insight
For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a BytesInput and passing it to PageWriter.writePage().
For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).
Proposed Design
A new ArrowParquetWriter in the parquet-arrow module that writes pages directly to PageWriter rather than going through RecordConsumer/ColumnWriter:
ArrowParquetWriter.writeBatch(VectorSchemaRoot)
→ per column: selects optimal write strategy
→ writes pages to PageWriter via getPageWriter(ColumnDescriptor)
→ manages row groups via ParquetFileWriter
Per-column strategy selection (best to worst):
- Zero-copy (non-null + fixed-width + PLAIN): wrap Arrow buffer as
BytesInputdirectly. RL/DL as single-value RLE runs. Stats from sequential buffer scan. - Bulk-copy with nulls (nullable + fixed-width + PLAIN): scan validity bitmap for non-null runs, bulk-copy each, emit DL as RLE runs per null-transition.
- Variable-width rewrite (PLAIN + string/binary): single-pass transformation of Arrow offset+data buffers to Parquet's length-prefixed format.
- Dictionary: map Arrow's
DictionaryEncodedVectorto Parquet dictionary pages, or build dictionary from plain vector. - Fallback: per-value through
ValuesWriter(same cost as today, for unsupported encoding/type combos).
All verified public APIs exist for this:
PageWriteStore.getPageWriter(ColumnDescriptor)— get column page writer directlyPageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings)— write pre-encoded pagesBytesInput.from(ByteBuffer)— zero-copy buffer wrappingRunLengthBitPackingHybridEncoder— encode RL/DL levelsColumnChunkPageWriteStore.flushToFileWriter()— flush pages to file
Why Not Extend ParquetWriter?
ParquetWriter.write(T) calls InternalParquetRecordWriter.write() which increments recordCount by 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manage ParquetFileWriter and row groups directly.
Scope
- Parquet file format unchanged
- Existing
ParquetWriter<T>API unchanged parquet-arrowgainsparquet-hadoopas a compile dependency (forParquetFileWriter,ColumnChunkPageWriteStore)- No Hadoop runtime dependency for uncompressed output (compression requires caller-supplied
CompressionCodecFactory) - Flat schemas (primitive columns) first; nested types deferred
Implementation Phases
- Core infrastructure + zero-copy for non-null fixed-width PLAIN columns
- Nullable column support (validity bitmap scanning + run-based bulk copy)
- Variable-width types (string/binary offset rewriting)
- Dictionary encoding support
- Bloom filter, nested types, advanced features
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 4d 12h
- Merged PRs (30d)
- 28
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/parquet-java
-
Type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
apache/parquet-java#3792 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
apache/parquet-java#3767 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
apache/parquet-java#3695 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
apache/parquet-java#3667 ·
-
Type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
apache/parquet-java#3587 ·
All issues in apache/parquet-java
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
elastic/gradle-plugins#157 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
cryptomator/hub#497 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
johanhaleby/occurrent#1120 ·