Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Add batch/columnar write API for Arrow VectorSchemaRoot (Java parity with C++/Python)

未关闭
#3,733 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
30/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
活跃
技术栈
java

调研方向

首先阅读 parquet-arrow 模块以及名为 PageWriteStorePageWriterBytesInputColumnChunkPageWriteStore 的 API。拟议的工作分为五个实现阶段,首先为非 nullable 的固定宽度 PLAIN 列实现零拷贝写入;完成情况应根据所选阶段进行评判,完整范围包括对 nullable、可变宽度和字典的支持。

由索引模型根据 Issue 内容生成。

描述

Problem

parquet-java has no public API to write Arrow VectorSchemaRoot to Parquet files. Callers must construct row objects and feed them through ParquetWriter<T>.write(T) one at a time. Arrow C++ and PyArrow support this natively via parquet::WriteTable() / pq.write_table().

Related prior discussion:

  • #2264 (open since 2019)
  • #3353
Motivation

Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy ParquetWriter's input API:

  • Apache Iceberg (FileAppender<Record>) — iceberg#17748
  • Apache Fluss (Arrow-native streaming storage) — fluss#4047
  • Apache Paimon (worked around this by building their own paimon-arrow writer that bypasses parquet-java's row API entirely)
Key Insight

For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a BytesInput and passing it to PageWriter.writePage().

For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).

Proposed Design

A new ArrowParquetWriter in the parquet-arrow module that writes pages directly to PageWriter rather than going through RecordConsumer/ColumnWriter:

ArrowParquetWriter.writeBatch(VectorSchemaRoot)
    → per column: selects optimal write strategy
    → writes pages to PageWriter via getPageWriter(ColumnDescriptor)
    → manages row groups via ParquetFileWriter

Per-column strategy selection (best to worst):

  1. Zero-copy (non-null + fixed-width + PLAIN): wrap Arrow buffer as BytesInput directly. RL/DL as single-value RLE runs. Stats from sequential buffer scan.
  2. Bulk-copy with nulls (nullable + fixed-width + PLAIN): scan validity bitmap for non-null runs, bulk-copy each, emit DL as RLE runs per null-transition.
  3. Variable-width rewrite (PLAIN + string/binary): single-pass transformation of Arrow offset+data buffers to Parquet's length-prefixed format.
  4. Dictionary: map Arrow's DictionaryEncodedVector to Parquet dictionary pages, or build dictionary from plain vector.
  5. Fallback: per-value through ValuesWriter (same cost as today, for unsupported encoding/type combos).

All verified public APIs exist for this:

  • PageWriteStore.getPageWriter(ColumnDescriptor) — get column page writer directly
  • PageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings) — write pre-encoded pages
  • BytesInput.from(ByteBuffer) — zero-copy buffer wrapping
  • RunLengthBitPackingHybridEncoder — encode RL/DL levels
  • ColumnChunkPageWriteStore.flushToFileWriter() — flush pages to file
Why Not Extend ParquetWriter?

ParquetWriter.write(T) calls InternalParquetRecordWriter.write() which increments recordCount by 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manage ParquetFileWriter and row groups directly.

Scope
  • Parquet file format unchanged
  • Existing ParquetWriter<T> API unchanged
  • parquet-arrow gains parquet-hadoop as a compile dependency (for ParquetFileWriter, ColumnChunkPageWriteStore)
  • No Hadoop runtime dependency for uncompressed output (compression requires caller-supplied CompressionCodecFactory)
  • Flat schemas (primitive columns) first; nested types deferred
Implementation Phases
  1. Core infrastructure + zero-copy for non-null fixed-width PLAIN columns
  2. Nullable column support (validity bitmap scanning + run-based bulk copy)
  3. Variable-width types (string/binary offset rewriting)
  4. Dictionary encoding support
  5. Bloom filter, nested types, advanced features
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
4 天 12 小时
30 天内合并 PR
28

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/parquet-java 的其他 Issue

查看 apache/parquet-java 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。