Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[C++][Python] IPC reader rejects a DictionaryEncoding without indexType, which the format allows

Đang mở Phù hợp với người mới
#51,779 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

@CaptainAni187 đang làm issue này rồi.

Từ ngày 5/10/2026.

  • #51807 của @CaptainAni187 — đang mở

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
68/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
cpp, python
Lĩnh vực
backend, data

Hướng nghiên cứu

Phần reject nằm trong cpp/src/arrow/ipc/metadata_internal.cc quanh DictionaryEncoding.indexType (CHECK_FLATBUFFERS_NOT_NULL). Theo format/Schema.fbs, trường bị thiếu nghĩa là signed int32. Mặc định thành int32() khi indexType là null thay vì thất bại, rồi thêm một bài kiểm tra đọc IPC C++ (issue gồm một stream chỉ schema và một bản repro PyArrow) để open_stream thành công với dictionary<values=string, indices=int32>.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Component: C++ Component: Python

[!NOTE]
I discovered this issue and wrote it up with help from Claude Opus 5.5.

Describe the bug, including details regarding any error messages, version, and platform.

The format allows DictionaryEncoding.indexType to be omitted. format/Schema.fbs says:

If this field is null, the indices must be signed int32.

The C++ IPC reader instead requires the field:

auto int_data = encoding->indexType();
CHECK_FLATBUFFERS_NOT_NULL(int_data, "DictionaryEncoding.indexType");

So PyArrow can't read an Arrow IPC stream whose dictionary-encoded field omits indexType:

OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
Reproduction

PyArrow and apache-arrow (JavaScript) both set indexType when they write, so this script embeds a stream without it. The reader fails on the schema, so the stream holds only a schema: one field, text, of type dictionary<values=string, indices=int32>, with indexType omitted. It was made from a stream that PyArrow wrote; the script that made it is below.

"""Read an Arrow IPC stream whose DictionaryEncoding omits indexType."""

import base64

import pyarrow as pa

# A schema-only stream with one field, text: dictionary-encoded strings, with
# DictionaryEncoding.indexType omitted, so the indices are int32 per Schema.fbs.
data = base64.b64decode(
    "/////5gAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAuP///wQAAAABAAAAFAAAABAAGAAI"
    "AAYABwAMABAAFAAQAAAAAAABBRQAAABEAAAAIAAAAAQAAAAAAAAABAAAAHRleHQAAAAACAAIAAAA"
    "BADc////DAAAAAgADAAIAAcACAAAAAAAAAEgAAAABAAEAAQAAAAIAAgAAAAAAP////8AAAAA"
)
print(pa.ipc.open_stream(data).schema)

Expected: text: dictionary<values=string, indices=int32, ordered=0>

Actual: OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata

apache-arrow (JavaScript) 21.2.0 reads the same stream as Dictionary<Int32, Utf8>, as the spec describes.

The error is the same in PyArrow 12.0.1, 16.1.0, 20.0.0, 24.0.0, and 25.0.1 (Python 3.11 and 3.14, macOS), so this isn't a regression.

How the stream was made

This script writes a schema-only stream with PyArrow, then omits indexType from its schema message. With PyArrow 25.0.1 it prints exactly the base64 above.

"""Write the stream above: a schema-only stream from PyArrow, with indexType omitted."""

import base64
import struct

import pyarrow as pa


def table_at(buf, pos):
    """Return the (table, vtable) positions for the table referenced at pos."""
    table = pos + struct.unpack_from("<I", buf, pos)[0]
    return table, table - struct.unpack_from("<i", buf, table)[0]


def field_offset(buf, table, vtable, slot):
    """Return a field's offset within its table, or 0 if the field is absent."""
    entry = 4 + 2 * slot
    present = entry < struct.unpack_from("<H", buf, vtable)[0]
    return struct.unpack_from("<H", buf, vtable + entry)[0] if present else 0


def without_index_type(stream):
    """Omit DictionaryEncoding.indexType from the first field of the schema message.

    Flatbuffers share identical vtables between tables (here the Schema's and the
    DictionaryEncoding's), so the encoding gets its own copy of its vtable, appended
    to the message, with indexType cleared.
    """
    buf = bytearray(stream)
    length = struct.unpack_from("<i", buf, 4)[0]  # After the 0xFFFFFFFF continuation marker
    meta = 8
    message, message_vt = table_at(buf, meta)
    schema, schema_vt = table_at(buf, message + field_offset(buf, message, message_vt, 2))  # Message.header
    fields = schema + field_offset(buf, schema, schema_vt, 1)  # Schema.fields
    field, field_vt = table_at(buf, fields + struct.unpack_from("<I", buf, fields)[0] + 4)  # fields[0]
    encoding, encoding_vt = table_at(buf, field + field_offset(buf, field, field_vt, 4))  # Field.dictionary
    vtable = bytearray(buf[encoding_vt:encoding_vt + struct.unpack_from("<H", buf, encoding_vt)[0]])
    struct.pack_into("<H", vtable, 4 + 2 * 1, 0)  # DictionaryEncoding.indexType: absent
    struct.pack_into("<i", buf, encoding, encoding - (meta + length))
    metadata = bytes(buf[meta:meta + length]) + bytes(vtable)
    metadata += bytes(-len(metadata) % 8)
    return struct.pack("<Ii", 0xFFFFFFFF, len(metadata)) + metadata + bytes(buf[meta + length:])


schema = pa.schema({"text": pa.dictionary(pa.int32(), pa.string())})
sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, schema):
    pass
print(base64.b64encode(without_index_type(sink.getvalue().to_pybytes())).decode())
Suggested fix

Default to int32 when the field is absent:

std::shared_ptr<DataType> index_type;
auto int_data = encoding->indexType();
if (int_data == nullptr) {
  // Format/Schema.fbs: "If this field is null, the indices must be signed int32."
  index_type = int32();
} else {
  RETURN_NOT_OK(IntFromFlatbuffer(int_data, &index_type));
}
Related
Ngôn ngữ chính
C++
Star
17.2k
Fork
4.3k
Merge trung bình
4 ngày 8 giờ
Pull request đã merge (30 ngày)
99

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của apache/arrow

Tất cả issue của apache/arrow

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.