Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Add batch/columnar write API for Arrow VectorSchemaRoot (Java parity with C++/Python)

Open
#3,733 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
30/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
java

Research direction

Start by reading the parquet-arrow module and the named PageWriteStore, PageWriter, BytesInput, and ColumnChunkPageWriteStore APIs. The proposed work spans five implementation phases, beginning with zero-copy writes for non-null fixed-width PLAIN columns; completion should be judged against the chosen phase, with the full scope including nullable, variable-width, and dictionary support.

Written by the indexing model from the issue text.

Description

Problem

parquet-java has no public API to write Arrow VectorSchemaRoot to Parquet files. Callers must construct row objects and feed them through ParquetWriter<T>.write(T) one at a time. Arrow C++ and PyArrow support this natively via parquet::WriteTable() / pq.write_table().

Related prior discussion:

  • #2264 (open since 2019)
  • #3353
Motivation

Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy ParquetWriter's input API:

  • Apache Iceberg (FileAppender<Record>) — iceberg#17748
  • Apache Fluss (Arrow-native streaming storage) — fluss#4047
  • Apache Paimon (worked around this by building their own paimon-arrow writer that bypasses parquet-java's row API entirely)
Key Insight

For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a BytesInput and passing it to PageWriter.writePage().

For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).

Proposed Design

A new ArrowParquetWriter in the parquet-arrow module that writes pages directly to PageWriter rather than going through RecordConsumer/ColumnWriter:

ArrowParquetWriter.writeBatch(VectorSchemaRoot)
    → per column: selects optimal write strategy
    → writes pages to PageWriter via getPageWriter(ColumnDescriptor)
    → manages row groups via ParquetFileWriter

Per-column strategy selection (best to worst):

  1. Zero-copy (non-null + fixed-width + PLAIN): wrap Arrow buffer as BytesInput directly. RL/DL as single-value RLE runs. Stats from sequential buffer scan.
  2. Bulk-copy with nulls (nullable + fixed-width + PLAIN): scan validity bitmap for non-null runs, bulk-copy each, emit DL as RLE runs per null-transition.
  3. Variable-width rewrite (PLAIN + string/binary): single-pass transformation of Arrow offset+data buffers to Parquet's length-prefixed format.
  4. Dictionary: map Arrow's DictionaryEncodedVector to Parquet dictionary pages, or build dictionary from plain vector.
  5. Fallback: per-value through ValuesWriter (same cost as today, for unsupported encoding/type combos).

All verified public APIs exist for this:

  • PageWriteStore.getPageWriter(ColumnDescriptor) — get column page writer directly
  • PageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings) — write pre-encoded pages
  • BytesInput.from(ByteBuffer) — zero-copy buffer wrapping
  • RunLengthBitPackingHybridEncoder — encode RL/DL levels
  • ColumnChunkPageWriteStore.flushToFileWriter() — flush pages to file
Why Not Extend ParquetWriter?

ParquetWriter.write(T) calls InternalParquetRecordWriter.write() which increments recordCount by 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manage ParquetFileWriter and row groups directly.

Scope
  • Parquet file format unchanged
  • Existing ParquetWriter<T> API unchanged
  • parquet-arrow gains parquet-hadoop as a compile dependency (for ParquetFileWriter, ColumnChunkPageWriteStore)
  • No Hadoop runtime dependency for uncompressed output (compression requires caller-supplied CompressionCodecFactory)
  • Flat schemas (primitive columns) first; nested types deferred
Implementation Phases
  1. Core infrastructure + zero-copy for non-null fixed-width PLAIN columns
  2. Nullable column support (validity bitmap scanning + run-based bulk copy)
  3. Variable-width types (string/binary offset rewriting)
  4. Dictionary encoding support
  5. Bloom filter, nested types, advanced features
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
4d 12h
Merged PRs (30d)
28

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/parquet-java

All issues in apache/parquet-java

Similar issues

More Java issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.