Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[C++][Python] IPC reader rejects a DictionaryEncoding without indexType, which the format allows

未关闭 适合新手
#51,779 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

@CaptainAni187 已经在做这个了。

开始于 2026年10月5日。

  • #51807 来自 @CaptainAni187 —— 未关闭

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
68/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
cpp, python
领域
backend, data

调研方向

reject 位于 cpp/src/arrow/ipc/metadata_internal.cc 中 DictionaryEncoding.indexType 附近(CHECK_FLATBUFFERS_NOT_NULL)。根据 format/Schema.fbs,缺失字段表示 signed int32。当 indexType 为 null 时默认使用 int32() 而不是失败,然后添加一个 C++ IPC 读取测试(该 issue 包含仅 schema 的流和 PyArrow repro),使 open_stream 在 dictionary<values=string, indices=int32> 下成功。

由索引模型根据 Issue 内容生成。

描述

Component: C++ Component: Python

[!NOTE]
I discovered this issue and wrote it up with help from Claude Opus 5.5.

Describe the bug, including details regarding any error messages, version, and platform.

The format allows DictionaryEncoding.indexType to be omitted. format/Schema.fbs says:

If this field is null, the indices must be signed int32.

The C++ IPC reader instead requires the field:

auto int_data = encoding->indexType();
CHECK_FLATBUFFERS_NOT_NULL(int_data, "DictionaryEncoding.indexType");

So PyArrow can't read an Arrow IPC stream whose dictionary-encoded field omits indexType:

OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
Reproduction

PyArrow and apache-arrow (JavaScript) both set indexType when they write, so this script embeds a stream without it. The reader fails on the schema, so the stream holds only a schema: one field, text, of type dictionary<values=string, indices=int32>, with indexType omitted. It was made from a stream that PyArrow wrote; the script that made it is below.

"""Read an Arrow IPC stream whose DictionaryEncoding omits indexType."""

import base64

import pyarrow as pa

# A schema-only stream with one field, text: dictionary-encoded strings, with
# DictionaryEncoding.indexType omitted, so the indices are int32 per Schema.fbs.
data = base64.b64decode(
    "/////5gAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAuP///wQAAAABAAAAFAAAABAAGAAI"
    "AAYABwAMABAAFAAQAAAAAAABBRQAAABEAAAAIAAAAAQAAAAAAAAABAAAAHRleHQAAAAACAAIAAAA"
    "BADc////DAAAAAgADAAIAAcACAAAAAAAAAEgAAAABAAEAAQAAAAIAAgAAAAAAP////8AAAAA"
)
print(pa.ipc.open_stream(data).schema)

Expected: text: dictionary<values=string, indices=int32, ordered=0>

Actual: OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata

apache-arrow (JavaScript) 21.2.0 reads the same stream as Dictionary<Int32, Utf8>, as the spec describes.

The error is the same in PyArrow 12.0.1, 16.1.0, 20.0.0, 24.0.0, and 25.0.1 (Python 3.11 and 3.14, macOS), so this isn't a regression.

How the stream was made

This script writes a schema-only stream with PyArrow, then omits indexType from its schema message. With PyArrow 25.0.1 it prints exactly the base64 above.

"""Write the stream above: a schema-only stream from PyArrow, with indexType omitted."""

import base64
import struct

import pyarrow as pa


def table_at(buf, pos):
    """Return the (table, vtable) positions for the table referenced at pos."""
    table = pos + struct.unpack_from("<I", buf, pos)[0]
    return table, table - struct.unpack_from("<i", buf, table)[0]


def field_offset(buf, table, vtable, slot):
    """Return a field's offset within its table, or 0 if the field is absent."""
    entry = 4 + 2 * slot
    present = entry < struct.unpack_from("<H", buf, vtable)[0]
    return struct.unpack_from("<H", buf, vtable + entry)[0] if present else 0


def without_index_type(stream):
    """Omit DictionaryEncoding.indexType from the first field of the schema message.

    Flatbuffers share identical vtables between tables (here the Schema's and the
    DictionaryEncoding's), so the encoding gets its own copy of its vtable, appended
    to the message, with indexType cleared.
    """
    buf = bytearray(stream)
    length = struct.unpack_from("<i", buf, 4)[0]  # After the 0xFFFFFFFF continuation marker
    meta = 8
    message, message_vt = table_at(buf, meta)
    schema, schema_vt = table_at(buf, message + field_offset(buf, message, message_vt, 2))  # Message.header
    fields = schema + field_offset(buf, schema, schema_vt, 1)  # Schema.fields
    field, field_vt = table_at(buf, fields + struct.unpack_from("<I", buf, fields)[0] + 4)  # fields[0]
    encoding, encoding_vt = table_at(buf, field + field_offset(buf, field, field_vt, 4))  # Field.dictionary
    vtable = bytearray(buf[encoding_vt:encoding_vt + struct.unpack_from("<H", buf, encoding_vt)[0]])
    struct.pack_into("<H", vtable, 4 + 2 * 1, 0)  # DictionaryEncoding.indexType: absent
    struct.pack_into("<i", buf, encoding, encoding - (meta + length))
    metadata = bytes(buf[meta:meta + length]) + bytes(vtable)
    metadata += bytes(-len(metadata) % 8)
    return struct.pack("<Ii", 0xFFFFFFFF, len(metadata)) + metadata + bytes(buf[meta + length:])


schema = pa.schema({"text": pa.dictionary(pa.int32(), pa.string())})
sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, schema):
    pass
print(base64.b64encode(without_index_type(sink.getvalue().to_pybytes())).decode())
Suggested fix

Default to int32 when the field is absent:

std::shared_ptr<DataType> index_type;
auto int_data = encoding->indexType();
if (int_data == nullptr) {
  // Format/Schema.fbs: "If this field is null, the indices must be signed int32."
  index_type = int32();
} else {
  RETURN_NOT_OK(IntFromFlatbuffer(int_data, &index_type));
}
Related
主要语言
C++
星标
17.2k
派生
4.3k
平均合并
4 天 8 小时
30 天内合并 PR
99

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/arrow 的其他 Issue

查看 apache/arrow 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。