[Java][IPC] AbstractCompressionCodec.compress() writes prefix=0 for empty buffers, incompatible with C++/Python readers

未关闭 适合新手
#1,196 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
68/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
冷清
技术栈
java

调研方向

从 vector/src/main/java/org/apache/arrow/vector/compression/AbstractCompressionCodec.java 中的空缓冲区分支开始,将其前缀处理与文档中记载的 -1 sentinel 进行比较。使用 LZ4_FRAME 或 ZSTD,通过一个仅包含空字符串的 string 列复现该情况,然后验证 C++ 和 Python reader 能够读取生成的 IPC 流且不会出现解压错误。

由索引模型根据 Issue 内容生成。

描述

Type: bug
Describe the bug

When a buffer has writerIndex == 0 (e.g., a string column where all values
are empty string ""), AbstractCompressionCodec.compress() writes an 8-byte
buffer with uncompressed_length = 0 as a "shortcut for empty buffer":

https://github.com/apache/arrow-java/blob/main/vector/src/main/java/org/apache/arrow/vector/compression/AbstractCompressionCodec.java#L32-L39

if (uncompressedBuffer.writerIndex() == 0L) {
    // shortcut for empty buffer
    compressedBuffer.setLong(0, 0);  // prefix = 0
    ...
}

This has been present since the initial implementation (ARROW-11899, 2021).

The Java decompress() handles this correctly (if size==0 return empty),
but C++ and Python Arrow readers do not recognize prefix=0. They attempt
to decompress 0 bytes of data, which fails:

  • C++ (Arrow 1.0.0 ~ latest): IOError: Lz4 compressed input contains less than one frame
  • PyArrow 21.0: same error
Reproduction

Write an Arrow IPC stream with LZ4_FRAME (or ZSTD) compression where one
string column has all values = ""
(empty string, not null). The string data
buffer has writerIndex = 0, triggering the empty buffer path.

// Writer
ArrowStreamWriter writer = new ArrowStreamWriter(root, null, channel,
    IpcOption.DEFAULT, CommonsCompressionFactory.INSTANCE, CodecType.LZ4_FRAME);

// All rows: stringVector.setSafe(i, "".getBytes());

Reading with C++ or Python fails at the first RecordBatch.

Root cause

The Arrow IPC compression format defines:

  • prefix > 0: compressed data follows, decompress to prefix bytes
  • prefix = -1: buffer stored uncompressed (sentinel)
  • prefix = 0: undefined — not in spec, not handled by C++/Python

Java writes prefix=0 for empty buffers, but only Java itself knows how to
read it back. C++/Python treat it as "0 bytes to decompress" → fail.

Suggested fix

Change the empty buffer path to use -1 sentinel (which all readers support):

if (uncompressedBuffer.writerIndex() == 0L) {
    ArrowBuf compressedBuffer = allocator.buffer(SIZE_OF_UNCOMPRESSED_LENGTH);
    compressedBuffer.setLong(0, -1L);  // Use -1 instead of 0
    compressedBuffer.writerIndex(SIZE_OF_UNCOMPRESSED_LENGTH);
    uncompressedBuffer.close();
    return compressedBuffer;
}

When a reader sees prefix=-1, it returns an empty/zero-length slice — which
is correct for an originally empty buffer.

Environment
  • Affected: All Arrow Java versions with IPC compression (1.0.0+)
  • Readers that fail: Arrow C++ (all versions), PyArrow (all versions)
  • Codec: Both LZ4_FRAME and ZSTD
Related
  • #1116 — similar symptom (prefix=0) but different root cause (race condition
    in vector reuse, not the intentional empty buffer path)
  • apache/arrow#15102 — C++ DecompressBuffer fix for prefix=-1 (does not
    handle prefix=0)
主要语言
Java
星标
95
派生
154
平均合并
2 天 16 小时
30 天内合并 PR
9

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/arrow-java 的其他 Issue

查看 apache/arrow-java 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。