Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Apache Parquet Java Performance Improvements

未关闭
#3,530 5 条评论 7 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 2 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
20/100
Issue 类型
重构
描述清晰度
需要澄清
活跃度
停滞
技术栈
java

调研方向

首先查看开放的 PR #3566、#3567、#3568 和 #3570,因为优化工作已拆分为相互独立的子系统。针对相关的 codec 或压缩领域,运行 issue 中的基准测试命令,然后使用 parquet-column、parquet-common 和 parquet-hadoop 模块的测试套件验证完成情况。

由索引模型根据 Issue 内容生成。

描述

Targeted performance optimizations to parquet-java's encoding, decoding, and compression layers. Each PR addresses an independent codec or subsystem, replacing stream-wrapper abstractions with direct ByteBuffer I/O, eliminating per-value allocations, and adding JMH benchmarks to validate the improvements.

All benchmarks: JDK 25.0.3 (Temurin), JMH 1.37, Linux x86_64, 100K values/iteration.

Pull requests

PR Status Scope Headline result
#3565 Merged PLAIN encoding/decoding Decode 2.6–13x, encode 2–3.8x
#3566 Open DICTIONARY encoding/decoding Encode 1.5–100x (high-card / long binary)
#3567 Open DELTA_BINARY_PACKED, DELTA_LENGTH_BYTE_ARRAY, DELTA_BYTE_ARRAY Decode +22–31%, encode +3–7%
#3568 Open RLE/Bit-Packing hybrid codec Decode +30%, encode +5%
#3569 Merged BYTE_STREAM_SPLIT Decode 1.4–5.1x, encode 1.6–6.9x
#3570 Open Compression (Snappy, ZSTD, LZ4, GZIP, LZO, Brotli) Isolated decompress +15–25% (Snappy/Zstd/Gzip), Snappy compress up to ~2.9x; enables LZ4/Brotli/LZO without native libs
#3571 Merged Row group flush buffer lifecycle Correct resource management (not a perf win)

PRs 1–6 and 8 are independent and can be reviewed/merged in parallel. Additional PRs for column I/O (par7) and level write batching (par9) will follow once their dependencies (#3565, #3568) land.

Common optimization patterns

  • InputStream → ByteBuffer: Direct ByteBuffer (LITTLE_ENDIAN) access for getInt()/getLong()/getFloat()/getDouble() JVM intrinsics instead of per-byte stream reads.
  • Buffer reuse: Allocate once, grow lazily — no fresh allocation per page or per run.
  • pack32/unpack32 fast paths: Batch 32 values per packer/unpacker call (4x fewer invocations).
  • Eliminate intermediate copies: Write directly to output buffers, bypass stream wrappers.

Benchmark highlights

PLAIN (#3565)
Benchmark Master Optimized Speedup
decodeInt 425 M ops/s 5,427 M ops/s 12.8x
decodeBoolean 639 1,642 2.6x
encodeInt 148 559 3.8x
encodeBoolean 850 1,692 2.0x
DICTIONARY (#3566)
Benchmark Master Optimized Speedup
encodeBinary LOW len=1000 1.5 M ops/s 148.3 M ops/s ~100x
encodeBinary LOW len=100 13.2 107.8 8.2x
encodeFlba HIGH len=12 6.3 15.4 2.4x
encodeInt HIGH_CARD 14.9 23.5 1.58x
DELTA (#3567)
Component Avg improvement
DELTA_BINARY_PACKED decode +27%
DELTA_BYTE_ARRAY decode +31%
Long delta decode (TIMESTAMP_MILLIS pattern) +28%
Encoding (all delta variants) +3–7%
RLE (#3568)
Category Avg improvement
Direct decoder (packed data) +30%
Boolean decode (packed patterns) +14%
Encoder +5%
BYTE_STREAM_SPLIT (#3569)
Benchmark Master Optimized Speedup
decodeInt 203 M ops/s 1,034 M ops/s 5.1x
encodeDouble 53 365 6.9x
encodeLong 52 356 6.9x
encodeInt 99 515 5.2x
Compression (#3570)

Isolated CompressionBenchmark, ours/master geomean over 16 realistic encoding shapes (>1 = faster):

Codec compress decompress
Snappy 1.38x 1.25x
Zstd 1.00x 1.18x
Gzip 1.00x 1.15x
LZ4_RAW 0.96x 1.06x

Snappy peaks near 2.9x compress / 1.8x decompress on small, highly compressible pages. Compression is ~1–10% of end-to-end write for the fast codecs (up to ~46% for Gzip), so file-level impact scales with codec weight. Also migrates LZO from GPL to Apache 2.0 (aircompressor) and Brotli from abandoned jbrotli to brotli4j (adds aarch64 support); LZ4/Brotli/LZO now work without native Hadoop codecs.

Row group flush (#3571)

Correct resource management — releases column buffers during flush rather than after. Peak memory unchanged (peak occurs during page compression, not flush). Makes buffers GC-eligible sooner.

How to run the benchmarks

# Build the benchmark jar
./mvnw --projects parquet-benchmarks -amd -DskipTests -Denforcer.skip=true clean package

# Quick single-iteration smoke test
./parquet-benchmarks/run.sh all -wi 0 -i 1 -f 1

# Full statistical run
./parquet-benchmarks/run.sh all -wi 5 -i 5 -f 3 -rff /tmp/benchmark-results.json

# Specific benchmark
./parquet-benchmarks/run.sh all "PlainEncodingBenchmark|PlainDecodingBenchmark"

Test validation

Each PR passes the full module test suite:

./mvnw --projects parquet-column,parquet-common,parquet-hadoop -amd verify
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
4 天 5 小时
30 天内合并 PR
30

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 有 Pull Request 模板
  • 没有贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/parquet-java 的其他 Issue

查看 apache/parquet-java 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。