[C++][Python] IPC reader rejects a DictionaryEncoding without indexType, which the format allows
I maintainer di solito rispondono entro 1 giorno
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 68/100
Direzione di ricerca
Il reject è in cpp/src/arrow/ipc/metadata_internal.cc intorno a DictionaryEncoding.indexType (CHECK_FLATBUFFERS_NOT_NULL). Secondo format/Schema.fbs, un campo mancante significa signed int32. Usare int32() di default quando indexType è null invece di fallire, poi aggiungere un test di lettura IPC in C++ (l'issue include uno stream solo schema e un repro PyArrow) così che open_stream abbia successo con dictionary<values=string, indices=int32>.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
[!NOTE]
I discovered this issue and wrote it up with help from Claude Opus 5.5.
Describe the bug, including details regarding any error messages, version, and platform.
The format allows DictionaryEncoding.indexType to be omitted. format/Schema.fbs says:
If this field is null, the indices must be signed int32.
The C++ IPC reader instead requires the field:
auto int_data = encoding->indexType();
CHECK_FLATBUFFERS_NOT_NULL(int_data, "DictionaryEncoding.indexType");
So PyArrow can't read an Arrow IPC stream whose dictionary-encoded field omits indexType:
OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
Reproduction
PyArrow and apache-arrow (JavaScript) both set indexType when they write, so this script embeds a stream without it. The reader fails on the schema, so the stream holds only a schema: one field, text, of type dictionary<values=string, indices=int32>, with indexType omitted. It was made from a stream that PyArrow wrote; the script that made it is below.
"""Read an Arrow IPC stream whose DictionaryEncoding omits indexType."""
import base64
import pyarrow as pa
# A schema-only stream with one field, text: dictionary-encoded strings, with
# DictionaryEncoding.indexType omitted, so the indices are int32 per Schema.fbs.
data = base64.b64decode(
"/////5gAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAuP///wQAAAABAAAAFAAAABAAGAAI"
"AAYABwAMABAAFAAQAAAAAAABBRQAAABEAAAAIAAAAAQAAAAAAAAABAAAAHRleHQAAAAACAAIAAAA"
"BADc////DAAAAAgADAAIAAcACAAAAAAAAAEgAAAABAAEAAQAAAAIAAgAAAAAAP////8AAAAA"
)
print(pa.ipc.open_stream(data).schema)
Expected: text: dictionary<values=string, indices=int32, ordered=0>
Actual: OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata
apache-arrow (JavaScript) 21.2.0 reads the same stream as Dictionary<Int32, Utf8>, as the spec describes.
The error is the same in PyArrow 12.0.1, 16.1.0, 20.0.0, 24.0.0, and 25.0.1 (Python 3.11 and 3.14, macOS), so this isn't a regression.
How the stream was made
This script writes a schema-only stream with PyArrow, then omits indexType from its schema message. With PyArrow 25.0.1 it prints exactly the base64 above.
"""Write the stream above: a schema-only stream from PyArrow, with indexType omitted."""
import base64
import struct
import pyarrow as pa
def table_at(buf, pos):
"""Return the (table, vtable) positions for the table referenced at pos."""
table = pos + struct.unpack_from("<I", buf, pos)[0]
return table, table - struct.unpack_from("<i", buf, table)[0]
def field_offset(buf, table, vtable, slot):
"""Return a field's offset within its table, or 0 if the field is absent."""
entry = 4 + 2 * slot
present = entry < struct.unpack_from("<H", buf, vtable)[0]
return struct.unpack_from("<H", buf, vtable + entry)[0] if present else 0
def without_index_type(stream):
"""Omit DictionaryEncoding.indexType from the first field of the schema message.
Flatbuffers share identical vtables between tables (here the Schema's and the
DictionaryEncoding's), so the encoding gets its own copy of its vtable, appended
to the message, with indexType cleared.
"""
buf = bytearray(stream)
length = struct.unpack_from("<i", buf, 4)[0] # After the 0xFFFFFFFF continuation marker
meta = 8
message, message_vt = table_at(buf, meta)
schema, schema_vt = table_at(buf, message + field_offset(buf, message, message_vt, 2)) # Message.header
fields = schema + field_offset(buf, schema, schema_vt, 1) # Schema.fields
field, field_vt = table_at(buf, fields + struct.unpack_from("<I", buf, fields)[0] + 4) # fields[0]
encoding, encoding_vt = table_at(buf, field + field_offset(buf, field, field_vt, 4)) # Field.dictionary
vtable = bytearray(buf[encoding_vt:encoding_vt + struct.unpack_from("<H", buf, encoding_vt)[0]])
struct.pack_into("<H", vtable, 4 + 2 * 1, 0) # DictionaryEncoding.indexType: absent
struct.pack_into("<i", buf, encoding, encoding - (meta + length))
metadata = bytes(buf[meta:meta + length]) + bytes(vtable)
metadata += bytes(-len(metadata) % 8)
return struct.pack("<Ii", 0xFFFFFFFF, len(metadata)) + metadata + bytes(buf[meta + length:])
schema = pa.schema({"text": pa.dictionary(pa.int32(), pa.string())})
sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, schema):
pass
print(base64.b64encode(without_index_type(sink.getvalue().to_pybytes())).decode())
Suggested fix
Default to int32 when the field is absent:
std::shared_ptr<DataType> index_type;
auto int_data = encoding->indexType();
if (int_data == nullptr) {
// Format/Schema.fbs: "If this field is null, the indices must be signed int32."
index_type = int32();
} else {
RETURN_NOT_OK(IntFromFlatbuffer(int_data, &index_type));
}
Related
- I found this through an apache-arrow (JavaScript) writer bug that drops
indexType, among other problems; that's reported separately in as https://github.com/apache/arrow-js/issues/496. This issue is only about reading a stream that is valid under the spec. - nanoarrow crashes on a stream like this one; that's reported separately at https://github.com/apache/arrow-nanoarrow/issues/956.
- Lingua principale
- C++
- Stelle
- 17.2k
- Fork
- 4.3k
- Merge medio
- 4g 5h
- PR unite (30g)
- 101
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di apache/arrow
-
Component: Ruby Type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
I maintainer di solito rispondono entro 1 giorno
-
[C++] GetSchema labels its null check on each schema field as "DictionaryEncoding.indexType"Forse già presa Una pull request collegata a questa issue è aperta o già unita. ApertaComponent: C++
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
Component: R Type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
apache/arrow#51695 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
[C++][Parquet] Plaintext-footer files written with AES_GCM_CTR_V1 record AES_GCM_V1 as the encryption algorithm and cannot be readForse già presa @YusefSyed l’ha presa 4 giorni fa. ApertaType: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
[R] Expose ignore_extra_columns and pad_short_rows CSV parse optionsForse già presa Una pull request collegata a questa issue è aperta o già unita. ApertaComponent: R good-first-issue Type: enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di apache/arrow
Issue simili
-
bug iOS 🍎 ui/ux
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
MerginMaps/mobile#4744 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
kokkos/kokkos-kernels#3328 ·
I maintainer di solito rispondono entro 1 giorno
-
bug needs triage tcp
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
project-chip/connectedhomeip#74644 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno