Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[C++][Python] IPC reader rejects a DictionaryEncoding without indexType, which the format allows

Aperta Adatta ai principianti
#51,779 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

@CaptainAni187 ci sta già lavorando.

Dal 5/10/2026.

  • #51807 di @CaptainAni187 — aperta

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
68/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
cpp, python
Ambito
backend, data

Direzione di ricerca

Il reject è in cpp/src/arrow/ipc/metadata_internal.cc intorno a DictionaryEncoding.indexType (CHECK_FLATBUFFERS_NOT_NULL). Secondo format/Schema.fbs, un campo mancante significa signed int32. Usare int32() di default quando indexType è null invece di fallire, poi aggiungere un test di lettura IPC in C++ (l'issue include uno stream solo schema e un repro PyArrow) così che open_stream abbia successo con dictionary<values=string, indices=int32>.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Component: C++ Component: Python

[!NOTE]
I discovered this issue and wrote it up with help from Claude Opus 5.5.

Describe the bug, including details regarding any error messages, version, and platform.

The format allows DictionaryEncoding.indexType to be omitted. format/Schema.fbs says:

If this field is null, the indices must be signed int32.

The C++ IPC reader instead requires the field:

auto int_data = encoding->indexType();
CHECK_FLATBUFFERS_NOT_NULL(int_data, "DictionaryEncoding.indexType");

So PyArrow can't read an Arrow IPC stream whose dictionary-encoded field omits indexType:

OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
Reproduction

PyArrow and apache-arrow (JavaScript) both set indexType when they write, so this script embeds a stream without it. The reader fails on the schema, so the stream holds only a schema: one field, text, of type dictionary<values=string, indices=int32>, with indexType omitted. It was made from a stream that PyArrow wrote; the script that made it is below.

"""Read an Arrow IPC stream whose DictionaryEncoding omits indexType."""

import base64

import pyarrow as pa

# A schema-only stream with one field, text: dictionary-encoded strings, with
# DictionaryEncoding.indexType omitted, so the indices are int32 per Schema.fbs.
data = base64.b64decode(
    "/////5gAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAuP///wQAAAABAAAAFAAAABAAGAAI"
    "AAYABwAMABAAFAAQAAAAAAABBRQAAABEAAAAIAAAAAQAAAAAAAAABAAAAHRleHQAAAAACAAIAAAA"
    "BADc////DAAAAAgADAAIAAcACAAAAAAAAAEgAAAABAAEAAQAAAAIAAgAAAAAAP////8AAAAA"
)
print(pa.ipc.open_stream(data).schema)

Expected: text: dictionary<values=string, indices=int32, ordered=0>

Actual: OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata

apache-arrow (JavaScript) 21.2.0 reads the same stream as Dictionary<Int32, Utf8>, as the spec describes.

The error is the same in PyArrow 12.0.1, 16.1.0, 20.0.0, 24.0.0, and 25.0.1 (Python 3.11 and 3.14, macOS), so this isn't a regression.

How the stream was made

This script writes a schema-only stream with PyArrow, then omits indexType from its schema message. With PyArrow 25.0.1 it prints exactly the base64 above.

"""Write the stream above: a schema-only stream from PyArrow, with indexType omitted."""

import base64
import struct

import pyarrow as pa


def table_at(buf, pos):
    """Return the (table, vtable) positions for the table referenced at pos."""
    table = pos + struct.unpack_from("<I", buf, pos)[0]
    return table, table - struct.unpack_from("<i", buf, table)[0]


def field_offset(buf, table, vtable, slot):
    """Return a field's offset within its table, or 0 if the field is absent."""
    entry = 4 + 2 * slot
    present = entry < struct.unpack_from("<H", buf, vtable)[0]
    return struct.unpack_from("<H", buf, vtable + entry)[0] if present else 0


def without_index_type(stream):
    """Omit DictionaryEncoding.indexType from the first field of the schema message.

    Flatbuffers share identical vtables between tables (here the Schema's and the
    DictionaryEncoding's), so the encoding gets its own copy of its vtable, appended
    to the message, with indexType cleared.
    """
    buf = bytearray(stream)
    length = struct.unpack_from("<i", buf, 4)[0]  # After the 0xFFFFFFFF continuation marker
    meta = 8
    message, message_vt = table_at(buf, meta)
    schema, schema_vt = table_at(buf, message + field_offset(buf, message, message_vt, 2))  # Message.header
    fields = schema + field_offset(buf, schema, schema_vt, 1)  # Schema.fields
    field, field_vt = table_at(buf, fields + struct.unpack_from("<I", buf, fields)[0] + 4)  # fields[0]
    encoding, encoding_vt = table_at(buf, field + field_offset(buf, field, field_vt, 4))  # Field.dictionary
    vtable = bytearray(buf[encoding_vt:encoding_vt + struct.unpack_from("<H", buf, encoding_vt)[0]])
    struct.pack_into("<H", vtable, 4 + 2 * 1, 0)  # DictionaryEncoding.indexType: absent
    struct.pack_into("<i", buf, encoding, encoding - (meta + length))
    metadata = bytes(buf[meta:meta + length]) + bytes(vtable)
    metadata += bytes(-len(metadata) % 8)
    return struct.pack("<Ii", 0xFFFFFFFF, len(metadata)) + metadata + bytes(buf[meta + length:])


schema = pa.schema({"text": pa.dictionary(pa.int32(), pa.string())})
sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, schema):
    pass
print(base64.b64encode(without_index_type(sink.getvalue().to_pybytes())).decode())
Suggested fix

Default to int32 when the field is absent:

std::shared_ptr<DataType> index_type;
auto int_data = encoding->indexType();
if (int_data == nullptr) {
  // Format/Schema.fbs: "If this field is null, the indices must be signed int32."
  index_type = int32();
} else {
  RETURN_NOT_OK(IntFromFlatbuffer(int_data, &index_type));
}
Related
Lingua principale
C++
Stelle
17.2k
Fork
4.3k
Merge medio
4g 5h
PR unite (30g)
101

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/arrow

Tutte le issue di apache/arrow

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.