[C++][Python] IPC reader rejects a DictionaryEncoding without indexType, which the format allows
维护者通常 1 天内回复
评估
调研方向
reject 位于 cpp/src/arrow/ipc/metadata_internal.cc 中 DictionaryEncoding.indexType 附近(CHECK_FLATBUFFERS_NOT_NULL)。根据 format/Schema.fbs,缺失字段表示 signed int32。当 indexType 为 null 时默认使用 int32() 而不是失败,然后添加一个 C++ IPC 读取测试(该 issue 包含仅 schema 的流和 PyArrow repro),使 open_stream 在 dictionary<values=string, indices=int32> 下成功。
由索引模型根据 Issue 内容生成。
描述
[!NOTE]
I discovered this issue and wrote it up with help from Claude Opus 5.5.
Describe the bug, including details regarding any error messages, version, and platform.
The format allows DictionaryEncoding.indexType to be omitted. format/Schema.fbs says:
If this field is null, the indices must be signed int32.
The C++ IPC reader instead requires the field:
auto int_data = encoding->indexType();
CHECK_FLATBUFFERS_NOT_NULL(int_data, "DictionaryEncoding.indexType");
So PyArrow can't read an Arrow IPC stream whose dictionary-encoded field omits indexType:
OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
Reproduction
PyArrow and apache-arrow (JavaScript) both set indexType when they write, so this script embeds a stream without it. The reader fails on the schema, so the stream holds only a schema: one field, text, of type dictionary<values=string, indices=int32>, with indexType omitted. It was made from a stream that PyArrow wrote; the script that made it is below.
"""Read an Arrow IPC stream whose DictionaryEncoding omits indexType."""
import base64
import pyarrow as pa
# A schema-only stream with one field, text: dictionary-encoded strings, with
# DictionaryEncoding.indexType omitted, so the indices are int32 per Schema.fbs.
data = base64.b64decode(
"/////5gAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAuP///wQAAAABAAAAFAAAABAAGAAI"
"AAYABwAMABAAFAAQAAAAAAABBRQAAABEAAAAIAAAAAQAAAAAAAAABAAAAHRleHQAAAAACAAIAAAA"
"BADc////DAAAAAgADAAIAAcACAAAAAAAAAEgAAAABAAEAAQAAAAIAAgAAAAAAP////8AAAAA"
)
print(pa.ipc.open_stream(data).schema)
Expected: text: dictionary<values=string, indices=int32, ordered=0>
Actual: OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
apache-arrow (JavaScript) 21.2.0 reads the same stream as Dictionary<Int32, Utf8>, as the spec describes.
The error is the same in PyArrow 12.0.1, 16.1.0, 20.0.0, 24.0.0, and 25.0.1 (Python 3.11 and 3.14, macOS), so this isn't a regression.
How the stream was made
This script writes a schema-only stream with PyArrow, then omits indexType from its schema message. With PyArrow 25.0.1 it prints exactly the base64 above.
"""Write the stream above: a schema-only stream from PyArrow, with indexType omitted."""
import base64
import struct
import pyarrow as pa
def table_at(buf, pos):
"""Return the (table, vtable) positions for the table referenced at pos."""
table = pos + struct.unpack_from("<I", buf, pos)[0]
return table, table - struct.unpack_from("<i", buf, table)[0]
def field_offset(buf, table, vtable, slot):
"""Return a field's offset within its table, or 0 if the field is absent."""
entry = 4 + 2 * slot
present = entry < struct.unpack_from("<H", buf, vtable)[0]
return struct.unpack_from("<H", buf, vtable + entry)[0] if present else 0
def without_index_type(stream):
"""Omit DictionaryEncoding.indexType from the first field of the schema message.
Flatbuffers share identical vtables between tables (here the Schema's and the
DictionaryEncoding's), so the encoding gets its own copy of its vtable, appended
to the message, with indexType cleared.
"""
buf = bytearray(stream)
length = struct.unpack_from("<i", buf, 4)[0] # After the 0xFFFFFFFF continuation marker
meta = 8
message, message_vt = table_at(buf, meta)
schema, schema_vt = table_at(buf, message + field_offset(buf, message, message_vt, 2)) # Message.header
fields = schema + field_offset(buf, schema, schema_vt, 1) # Schema.fields
field, field_vt = table_at(buf, fields + struct.unpack_from("<I", buf, fields)[0] + 4) # fields[0]
encoding, encoding_vt = table_at(buf, field + field_offset(buf, field, field_vt, 4)) # Field.dictionary
vtable = bytearray(buf[encoding_vt:encoding_vt + struct.unpack_from("<H", buf, encoding_vt)[0]])
struct.pack_into("<H", vtable, 4 + 2 * 1, 0) # DictionaryEncoding.indexType: absent
struct.pack_into("<i", buf, encoding, encoding - (meta + length))
metadata = bytes(buf[meta:meta + length]) + bytes(vtable)
metadata += bytes(-len(metadata) % 8)
return struct.pack("<Ii", 0xFFFFFFFF, len(metadata)) + metadata + bytes(buf[meta + length:])
schema = pa.schema({"text": pa.dictionary(pa.int32(), pa.string())})
sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, schema):
pass
print(base64.b64encode(without_index_type(sink.getvalue().to_pybytes())).decode())
Suggested fix
Default to int32 when the field is absent:
std::shared_ptr<DataType> index_type;
auto int_data = encoding->indexType();
if (int_data == nullptr) {
// Format/Schema.fbs: "If this field is null, the indices must be signed int32."
index_type = int32();
} else {
RETURN_NOT_OK(IntFromFlatbuffer(int_data, &index_type));
}
Related
- I found this through an apache-arrow (JavaScript) writer bug that drops
indexType, among other problems; that's reported separately in as https://github.com/apache/arrow-js/issues/496. This issue is only about reading a stream that is valid under the spec. - nanoarrow crashes on a stream like this one; that's reported separately at https://github.com/apache/arrow-nanoarrow/issues/956.
- 主要语言
- C++
- 星标
- 17.2k
- 派生
- 4.3k
- 平均合并
- 4 天 8 小时
- 30 天内合并 PR
- 99
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/arrow 的其他 Issue
-
[C++] GetSchema labels its null check on each schema field as "DictionaryEncoding.indexType"可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭Component: C++
难度 1/5 1 小时以内 新手友好度 88/100
维护者通常 1 天内回复
-
Component: R Type: bug
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
[C++][Parquet] Plaintext-footer files written with AES_GCM_CTR_V1 record AES_GCM_V1 as the encryption algorithm and cannot be read可能已有人在做 @YusefSyed 于 4 天前认领。 未关闭Type: bug
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复
-
[R] Expose ignore_extra_columns and pad_short_rows CSV parse options可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭Component: R good-first-issue Type: enhancement
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复
-
Component: C++
难度 2/5 1-3 小时 新手友好度 86/100
维护者通常 1 天内回复
相似的 Issue
-
bug
难度 2/5 1-3 小时 新手友好度 75/100
维护者通常 1 天内回复
-
lldb
难度 2/5 1-3 小时 新手友好度 85/100
llvm/llvm-project#229592 · 11 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 66/100
维护者通常 3 天内回复
-
难度 1/5 1 小时以内 新手友好度 94/100
llvm/offload-test-suite#1560 ·
维护者通常 1 天内回复
-
agent:Windows bug MEDIUM performance tooling
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复