[C++][Substrait] RelCommon.emit indices are not bounds-checked before they index the schema
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 新手友好度
- 72/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
调研方向
从 cpp/src/arrow/engine/substrait/relation_internal.cc 开始,阅读 GetEmitInfo、ProcessEmitProject 和 ProcessExtensionEmit,比较它们对 emit 索引的处理方式。运行所提供的 emit_repro.py 变体,以复现崩溃和无效索引行为。完成标准是:超出范围的映射返回错误,而不是索引到 schema 之外,同时有效映射仍能正常工作。
由索引模型根据 Issue 内容生成。
描述
GetEmitInfo passes every value in RelCommon.emit.output_mapping to both FieldRef(map_id) and the unchecked input_schema->field(map_id). ProcessEmitProject does the same on the project path, where the value also indexes proj_options.expressions. With this implementation, a positive out-of-range index still reads past the end of a FieldVector and the process dies. For -1, Acero later rejects FieldRef(-1) with ArrowInvalid, but the unchecked schema lookup happens first. ProcessExtensionEmit, in the same file, already returns Status::Invalid("Out of bounds emit index ", emit_idx).
read, emit [1, 0] [DataType(double), DataType(int64)]
read, emit [5] exit 139
read, emit [-1] ArrowInvalid: No match for FieldRef.FieldPath(-1)
project, emit [5] exit 139
aggregate, emit [1, 0] exit 139
The last variant is why this is more than a check on malformed input: that plan is valid Substrait. Arrow's vendored proto has no expression_references, so Arrow drops the grouping keys and derives an empty aggregate schema. The plan's valid [1, 0] mapping is then out of range for that empty schema. That grouping-key loss is #50634. A bounds check would turn the process crash into an error that reports the rejected [1, 0] mapping; fixing #50634 is still required for the plan to run.
Rechecked with PyArrow 26.0.0.dev244+g79e074ace, built from main at 79e074ace3a7c4ca26f211dfd84a1984ad53c362, on macOS arm64. The two positive out-of-range cases and the aggregate case still exit 139; the negative case now returns ArrowInvalid. The original PyArrow 25.0.1 reproduction crashed for the negative case as well on macOS arm64 and Linux x86-64. The source is cpp/src/arrow/engine/substrait/relation_internal.cc.
Reproducer — one file, pyarrow only
"""RelCommon.emit.output_mapping indices are used to index the input schema unchecked.
The control succeeds and the negative case is caught as `ArrowInvalid`. Run each variant in a process of its own because the other three terminate:
for v in 0 1 2 3 4; do python3 emit_repro.py $v || echo " exit $?"; done
"""
import json, sys
import pyarrow as pa
import pyarrow.substrait as ps
from pyarrow._substrait import _parse_json_plan
SCHEMA = pa.schema([pa.field("c0", pa.int64(), nullable=False),
pa.field("c1", pa.float64(), nullable=False)])
def provider(names, schema=None):
return pa.table({"c0": [1], "c1": [2.5]}, schema=schema or SCHEMA)
def ref(i):
return {"selection": {"directReference": {"structField": {"field": i} if i else {}},
"rootReference": {}}}
def emit(m):
return {"common": {"emit": {"outputMapping": m}}}
def read(m=None):
r = {"baseSchema": {"names": ["c0", "c1"], "struct": {
"types": [{"i64": {"nullability": "NULLABILITY_REQUIRED"}},
{"fp64": {"nullability": "NULLABILITY_REQUIRED"}}],
"nullability": "NULLABILITY_REQUIRED"}},
"namedTable": {"names": ["t"]}}
return {"read": dict(r, **(emit(m) if m else {}))}
def run(rel, names):
plan = {"version": {"minorNumber": 102, "producer": "repro"},
"relations": [{"root": {"input": rel, "names": names}}]}
return ps.run_query(_parse_json_plan(json.dumps(plan).encode()), table_provider=provider)
# A plan that is valid Substrait: the grouping keys are where current Substrait puts them.
# Arrow's vendored proto has no expression_references, so it drops them (#50634) and the
# aggregate's output schema has no fields at all - which puts a correct mapping out of range.
AGG = {"aggregate": dict({"input": read(),
"groupings": [{"expressionReferences": [0, 1]}],
"groupingExpressions": [ref(0), ref(1)]}, **emit([1, 0]))}
VARIANTS = [
("read, emit [1, 0]", lambda: run(read([1, 0]), ["c1", "c0"])),
("read, emit [5]", lambda: run(read([5]), ["x"])),
("read, emit [-1]", lambda: run(read([-1]), ["x"])),
("project, emit [5]", lambda: run({"project": dict(
{"input": read(), "expressions": [ref(0)]},
**emit([5]))}, ["x"])),
("aggregate, emit [1, 0]", lambda: run(AGG, ["c1", "c0"])),
]
label, case = VARIANTS[int(sys.argv[1])]
print("%-22s" % label, end=" ", flush=True)
try:
print(case().read_all().schema.types)
except Exception as e:
print(type(e).__name__ + ":", str(e).replace("\n", " ")[:70])
- 主要语言
- C++
- 星标
- 17.2k
- 派生
- 4.3k
- 平均合并
- 4 天 3 小时
- 30 天内合并 PR
- 93
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/arrow 的其他 Issue
-
Component: R Type: bug
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
[C++][Parquet] Plaintext-footer files written with AES_GCM_CTR_V1 record AES_GCM_V1 as the encryption algorithm and cannot be read可能已有人在做 @YusefSyed 于 3 天前认领。 未关闭Type: bug
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复
-
[R] Expose ignore_extra_columns and pad_short_rows CSV parse options可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭Component: R good-first-issue Type: enhancement
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复
-
Component: C++
难度 2/5 1-3 小时 新手友好度 86/100
维护者通常 1 天内回复
-
Component: C++ Type: enhancement
难度 2/5 1-3 小时 新手友好度 70/100
维护者通常 1 天内回复
相似的 Issue
-
code-quality libc++
难度 1/5 1 小时以内 新手友好度 82/100
llvm/llvm-project#229284 ·
维护者通常 1 天内回复
-
test-issue
难度 2/5 1-3 小时 新手友好度 82/100
llvm/offload-test-suite#1557 ·
维护者通常 1 天内回复
-
enhancement
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 85/100
maplibre/maplibre-native#4723 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
HarbourMasters/Shipwright#7320 ·
维护者通常 1 天内回复