Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Empty ListVector and LargeListVector can expose offset buffers with writerIndex greater than capacity

未关闭
#1,194 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
72/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
冷清
技术栈
java
领域
data

调研方向

首先检查 org.apache.arrow.vector.complex.ListVector 和 LargeListVector 中的 setReaderAndWriterIndex(),然后跟踪 valueCount == 0 时它们的偏移缓冲区是如何分配的。验证导出的缓冲区保留开头的零偏移,并且对于这两种向量类型都满足 writerIndex <= capacity,同时不缩小未来的偏移分配。

由索引模型根据 Issue 内容生成。

描述

Describe the bug, including details regarding any error messages, version, and platform.

ListVector and LargeListVector can expose an invalid offset buffer state when valueCount == 0.

For an empty list vector, the logical offset buffer should still contain the leading offset entry:

  • ListVector: (valueCount + 1) * 4 == 4 bytes
  • LargeListVector: (valueCount + 1) * 8 == 8 bytes

However, in the empty-vector path, the offset buffer can have:

readerIndex: 0
writerIndex: 4
capacity: 0

or the equivalent writerIndex: 8, capacity: 0 for LargeListVector.

This violates the normal buffer invariant:

0 <= readerIndex <= writerIndex <= capacity

Downstream consumers that unwrap or serialize the Arrow buffer through Netty can then fail with:

IndexOutOfBoundsException: readerIndex: 0, writerIndex: 4
(expected: 0 <= readerIndex <= writerIndex <= capacity(0))

The issue is that setReaderAndWriterIndex() sets the offset buffer writer index based on valueCount * OFFSET_WIDTH, which is 0 for empty vectors. But list vectors still require one offset slot even when there are no values.

The same issue applies to both:

  • org.apache.arrow.vector.complex.ListVector
  • org.apache.arrow.vector.complex.LargeListVector
Expected behavior

For valueCount == 0, the offset buffer should still have enough capacity and readable bytes for the leading zero offset:

(valueCount + 1) * OFFSET_WIDTH

So:

  • empty ListVector should expose at least 4 bytes for offset [0]
  • empty LargeListVector should expose at least 8 bytes for offset [0]

The first offset value should be zero.

Actual behavior

An empty list vector can expose an offset buffer with a non-zero writer index but zero capacity, causing Netty buffer validation to fail when the buffer is unwrapped or consumed.

Suggested fix

Update ListVector.setReaderAndWriterIndex() and LargeListVector.setReaderAndWriterIndex() so the offset buffer writer index is based on:

(valueCount + 1) * OFFSET_WIDTH

For the valueCount == 0 case, ensure the offset buffer has enough capacity for the leading zero offset before setting the writer index.

Care should be taken not to shrink the vector's future offset allocation size when allocating this empty sentinel offset buffer.

Additional context

This was observed downstream in Dremio after upgrading Arrow Java. The failure occurred while sending a record batch containing an empty list vector, where the send path unwraps Arrow buffers through Netty.

The downstream error was:

SYSTEM ERROR: IndexOutOfBoundsException: readerIndex: 0, writerIndex: 4
(expected: 0 <= readerIndex <= writerIndex <= capacity(0))

This issue is distinct from #1125. That issue involves UnionListReader.setPosition on a post-IPC empty list. This issue is about the offset buffer exported by empty ListVector / LargeListVector instances having an invalid writer-index/capacity relationship.

主要语言
Java
星标
95
派生
154
平均合并
2 天 10 小时
30 天内合并 PR
11

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/arrow-java 的其他 Issue

查看 apache/arrow-java 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。