[C++][Python] IPC reader rejects a DictionaryEncoding without indexType, which the format allows
Maintainer thường phản hồi trong vòng 1 ngày
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 68/100
Hướng nghiên cứu
Phần reject nằm trong cpp/src/arrow/ipc/metadata_internal.cc quanh DictionaryEncoding.indexType (CHECK_FLATBUFFERS_NOT_NULL). Theo format/Schema.fbs, trường bị thiếu nghĩa là signed int32. Mặc định thành int32() khi indexType là null thay vì thất bại, rồi thêm một bài kiểm tra đọc IPC C++ (issue gồm một stream chỉ schema và một bản repro PyArrow) để open_stream thành công với dictionary<values=string, indices=int32>.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
[!NOTE]
I discovered this issue and wrote it up with help from Claude Opus 5.5.
Describe the bug, including details regarding any error messages, version, and platform.
The format allows DictionaryEncoding.indexType to be omitted. format/Schema.fbs says:
If this field is null, the indices must be signed int32.
The C++ IPC reader instead requires the field:
auto int_data = encoding->indexType();
CHECK_FLATBUFFERS_NOT_NULL(int_data, "DictionaryEncoding.indexType");
So PyArrow can't read an Arrow IPC stream whose dictionary-encoded field omits indexType:
OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
Reproduction
PyArrow and apache-arrow (JavaScript) both set indexType when they write, so this script embeds a stream without it. The reader fails on the schema, so the stream holds only a schema: one field, text, of type dictionary<values=string, indices=int32>, with indexType omitted. It was made from a stream that PyArrow wrote; the script that made it is below.
"""Read an Arrow IPC stream whose DictionaryEncoding omits indexType."""
import base64
import pyarrow as pa
# A schema-only stream with one field, text: dictionary-encoded strings, with
# DictionaryEncoding.indexType omitted, so the indices are int32 per Schema.fbs.
data = base64.b64decode(
"/////5gAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAuP///wQAAAABAAAAFAAAABAAGAAI"
"AAYABwAMABAAFAAQAAAAAAABBRQAAABEAAAAIAAAAAQAAAAAAAAABAAAAHRleHQAAAAACAAIAAAA"
"BADc////DAAAAAgADAAIAAcACAAAAAAAAAEgAAAABAAEAAQAAAAIAAgAAAAAAP////8AAAAA"
)
print(pa.ipc.open_stream(data).schema)
Expected: text: dictionary<values=string, indices=int32, ordered=0>
Actual: OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
apache-arrow (JavaScript) 21.2.0 reads the same stream as Dictionary<Int32, Utf8>, as the spec describes.
The error is the same in PyArrow 12.0.1, 16.1.0, 20.0.0, 24.0.0, and 25.0.1 (Python 3.11 and 3.14, macOS), so this isn't a regression.
How the stream was made
This script writes a schema-only stream with PyArrow, then omits indexType from its schema message. With PyArrow 25.0.1 it prints exactly the base64 above.
"""Write the stream above: a schema-only stream from PyArrow, with indexType omitted."""
import base64
import struct
import pyarrow as pa
def table_at(buf, pos):
"""Return the (table, vtable) positions for the table referenced at pos."""
table = pos + struct.unpack_from("<I", buf, pos)[0]
return table, table - struct.unpack_from("<i", buf, table)[0]
def field_offset(buf, table, vtable, slot):
"""Return a field's offset within its table, or 0 if the field is absent."""
entry = 4 + 2 * slot
present = entry < struct.unpack_from("<H", buf, vtable)[0]
return struct.unpack_from("<H", buf, vtable + entry)[0] if present else 0
def without_index_type(stream):
"""Omit DictionaryEncoding.indexType from the first field of the schema message.
Flatbuffers share identical vtables between tables (here the Schema's and the
DictionaryEncoding's), so the encoding gets its own copy of its vtable, appended
to the message, with indexType cleared.
"""
buf = bytearray(stream)
length = struct.unpack_from("<i", buf, 4)[0] # After the 0xFFFFFFFF continuation marker
meta = 8
message, message_vt = table_at(buf, meta)
schema, schema_vt = table_at(buf, message + field_offset(buf, message, message_vt, 2)) # Message.header
fields = schema + field_offset(buf, schema, schema_vt, 1) # Schema.fields
field, field_vt = table_at(buf, fields + struct.unpack_from("<I", buf, fields)[0] + 4) # fields[0]
encoding, encoding_vt = table_at(buf, field + field_offset(buf, field, field_vt, 4)) # Field.dictionary
vtable = bytearray(buf[encoding_vt:encoding_vt + struct.unpack_from("<H", buf, encoding_vt)[0]])
struct.pack_into("<H", vtable, 4 + 2 * 1, 0) # DictionaryEncoding.indexType: absent
struct.pack_into("<i", buf, encoding, encoding - (meta + length))
metadata = bytes(buf[meta:meta + length]) + bytes(vtable)
metadata += bytes(-len(metadata) % 8)
return struct.pack("<Ii", 0xFFFFFFFF, len(metadata)) + metadata + bytes(buf[meta + length:])
schema = pa.schema({"text": pa.dictionary(pa.int32(), pa.string())})
sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, schema):
pass
print(base64.b64encode(without_index_type(sink.getvalue().to_pybytes())).decode())
Suggested fix
Default to int32 when the field is absent:
std::shared_ptr<DataType> index_type;
auto int_data = encoding->indexType();
if (int_data == nullptr) {
// Format/Schema.fbs: "If this field is null, the indices must be signed int32."
index_type = int32();
} else {
RETURN_NOT_OK(IntFromFlatbuffer(int_data, &index_type));
}
Related
- I found this through an apache-arrow (JavaScript) writer bug that drops
indexType, among other problems; that's reported separately in as https://github.com/apache/arrow-js/issues/496. This issue is only about reading a stream that is valid under the spec. - nanoarrow crashes on a stream like this one; that's reported separately at https://github.com/apache/arrow-nanoarrow/issues/956.
- Ngôn ngữ chính
- C++
- Star
- 17.2k
- Fork
- 4.3k
- Merge trung bình
- 4 ngày 8 giờ
- Pull request đã merge (30 ngày)
- 99
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/arrow
-
[C++] GetSchema labels its null check on each schema field as "DictionaryEncoding.indexType"Có thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mởComponent: C++
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Component: R Type: bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
apache/arrow#51695 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[C++][Parquet] Plaintext-footer files written with AES_GCM_CTR_V1 record AES_GCM_V1 as the encryption algorithm and cannot be readCó thể đã có người làm @YusefSyed đã nhận 4 ngày trước. Đang mởType: bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày
-
[R] Expose ignore_extra_columns and pad_short_rows CSV parse optionsCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mởComponent: R good-first-issue Type: enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Component: C++
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
Maintainer thường phản hồi trong vòng 1 ngày
Issue tương tự
-
area:runtime good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
WATonomous/wato_f1tenth#39 ·
-
[APP BUG]: Sorting by name after searching can bring up irrelevant resultsCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
shadps4-emu/shadps4-qtlauncher#465 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
duckdb/duckdb-excel#104 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
lxqt/lxqt-powermanagement#495 ·
-
`FakeBackendV2.run` fails with `NoiseError` on circuits with delays on qubits where T2 > 2·T1Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Qiskit/qiskit-aer#2466 ·