[C++][Substrait] RelCommon.emit indices are not bounds-checked before they index the schema
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 3/5
- Tiempo estimado
- 1-2 días
- Aptitud para principiantes
- 72/100
- Tipo de issue
- Error
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Área
- data-engineering
Línea de trabajo
Comience con cpp/src/arrow/engine/substrait/relation_internal.cc y lea GetEmitInfo, ProcessEmitProject y ProcessExtensionEmit para comparar cómo manejan los índices de emisión. Ejecute las variantes proporcionadas de emit_repro.py para reproducir los cierres inesperados y el comportamiento de índices no válidos. La tarea estará terminada cuando las asignaciones fuera de rango devuelvan un error en lugar de indexar más allá del esquema, mientras que las asignaciones válidas sigan funcionando.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
GetEmitInfo passes every value in RelCommon.emit.output_mapping to both FieldRef(map_id) and the unchecked input_schema->field(map_id). ProcessEmitProject does the same on the project path, where the value also indexes proj_options.expressions. With this implementation, a positive out-of-range index still reads past the end of a FieldVector and the process dies. For -1, Acero later rejects FieldRef(-1) with ArrowInvalid, but the unchecked schema lookup happens first. ProcessExtensionEmit, in the same file, already returns Status::Invalid("Out of bounds emit index ", emit_idx).
read, emit [1, 0] [DataType(double), DataType(int64)]
read, emit [5] exit 139
read, emit [-1] ArrowInvalid: No match for FieldRef.FieldPath(-1)
project, emit [5] exit 139
aggregate, emit [1, 0] exit 139
The last variant is why this is more than a check on malformed input: that plan is valid Substrait. Arrow's vendored proto has no expression_references, so Arrow drops the grouping keys and derives an empty aggregate schema. The plan's valid [1, 0] mapping is then out of range for that empty schema. That grouping-key loss is #50634. A bounds check would turn the process crash into an error that reports the rejected [1, 0] mapping; fixing #50634 is still required for the plan to run.
Rechecked with PyArrow 26.0.0.dev244+g79e074ace, built from main at 79e074ace3a7c4ca26f211dfd84a1984ad53c362, on macOS arm64. The two positive out-of-range cases and the aggregate case still exit 139; the negative case now returns ArrowInvalid. The original PyArrow 25.0.1 reproduction crashed for the negative case as well on macOS arm64 and Linux x86-64. The source is cpp/src/arrow/engine/substrait/relation_internal.cc.
Reproducer — one file, pyarrow only
"""RelCommon.emit.output_mapping indices are used to index the input schema unchecked.
The control succeeds and the negative case is caught as `ArrowInvalid`. Run each variant in a process of its own because the other three terminate:
for v in 0 1 2 3 4; do python3 emit_repro.py $v || echo " exit $?"; done
"""
import json, sys
import pyarrow as pa
import pyarrow.substrait as ps
from pyarrow._substrait import _parse_json_plan
SCHEMA = pa.schema([pa.field("c0", pa.int64(), nullable=False),
pa.field("c1", pa.float64(), nullable=False)])
def provider(names, schema=None):
return pa.table({"c0": [1], "c1": [2.5]}, schema=schema or SCHEMA)
def ref(i):
return {"selection": {"directReference": {"structField": {"field": i} if i else {}},
"rootReference": {}}}
def emit(m):
return {"common": {"emit": {"outputMapping": m}}}
def read(m=None):
r = {"baseSchema": {"names": ["c0", "c1"], "struct": {
"types": [{"i64": {"nullability": "NULLABILITY_REQUIRED"}},
{"fp64": {"nullability": "NULLABILITY_REQUIRED"}}],
"nullability": "NULLABILITY_REQUIRED"}},
"namedTable": {"names": ["t"]}}
return {"read": dict(r, **(emit(m) if m else {}))}
def run(rel, names):
plan = {"version": {"minorNumber": 102, "producer": "repro"},
"relations": [{"root": {"input": rel, "names": names}}]}
return ps.run_query(_parse_json_plan(json.dumps(plan).encode()), table_provider=provider)
# A plan that is valid Substrait: the grouping keys are where current Substrait puts them.
# Arrow's vendored proto has no expression_references, so it drops them (#50634) and the
# aggregate's output schema has no fields at all - which puts a correct mapping out of range.
AGG = {"aggregate": dict({"input": read(),
"groupings": [{"expressionReferences": [0, 1]}],
"groupingExpressions": [ref(0), ref(1)]}, **emit([1, 0]))}
VARIANTS = [
("read, emit [1, 0]", lambda: run(read([1, 0]), ["c1", "c0"])),
("read, emit [5]", lambda: run(read([5]), ["x"])),
("read, emit [-1]", lambda: run(read([-1]), ["x"])),
("project, emit [5]", lambda: run({"project": dict(
{"input": read(), "expressions": [ref(0)]},
**emit([5]))}, ["x"])),
("aggregate, emit [1, 0]", lambda: run(AGG, ["c1", "c0"])),
]
label, case = VARIANTS[int(sys.argv[1])]
print("%-22s" % label, end=" ", flush=True)
try:
print(case().read_all().schema.types)
except Exception as e:
print(type(e).__name__ + ":", str(e).replace("\n", " ")[:70])
- Lenguaje dominante
- C++
- Estrellas
- 17.2k
- Forks
- 4.3k
- Merge medio
- 4 d 3 h
- PR fusionados (30 d)
- 93
Preparar el entorno
- Incluye un Dockerfile o un archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de apache/arrow
-
Component: R Type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
apache/arrow#51695 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
[C++][Parquet] Plaintext-footer files written with AES_GCM_CTR_V1 record AES_GCM_V1 as the encryption algorithm and cannot be readPosiblemente ocupada @YusefSyed la tomó hace 2 días. AbiertoType: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
[R] Expose ignore_extra_columns and pad_short_rows CSV parse optionsPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. AbiertoComponent: R good-first-issue Type: enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Los mantenedores suelen responder en 1 día
-
Component: C++
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Los mantenedores suelen responder en 1 día
-
Component: C++ Type: enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
Todos los issues de apache/arrow
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Icinga/icinga2#11077 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100
-
agent:WSL bug linux LOW ui
Dificultad 1/5 Menos de una hora Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Copter: PosHold brake-entry threshold became 16 deg instead of 0.16 deg after the radians conversionAbierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 78/100
ArduPilot/ardupilot#34617 · 1 comentario · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
tesseract-robotics/tesseract_nanobind#168 ·
Los mantenedores suelen responder en 1 día