Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[C++][Substrait] RelCommon.emit indices are not bounds-checked before they index the schema

未关闭
#51,243 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
72/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
cpp, python

调研方向

从 cpp/src/arrow/engine/substrait/relation_internal.cc 开始,阅读 GetEmitInfo、ProcessEmitProject 和 ProcessExtensionEmit,比较它们对 emit 索引的处理方式。运行所提供的 emit_repro.py 变体,以复现崩溃和无效索引行为。完成标准是:超出范围的映射返回错误,而不是索引到 schema 之外,同时有效映射仍能正常工作。

由索引模型根据 Issue 内容生成。

描述

GetEmitInfo passes every value in RelCommon.emit.output_mapping to both FieldRef(map_id) and the unchecked input_schema->field(map_id). ProcessEmitProject does the same on the project path, where the value also indexes proj_options.expressions. With this implementation, a positive out-of-range index still reads past the end of a FieldVector and the process dies. For -1, Acero later rejects FieldRef(-1) with ArrowInvalid, but the unchecked schema lookup happens first. ProcessExtensionEmit, in the same file, already returns Status::Invalid("Out of bounds emit index ", emit_idx).

read, emit [1, 0]      [DataType(double), DataType(int64)]
read, emit [5]            exit 139
read, emit [-1]           ArrowInvalid: No match for FieldRef.FieldPath(-1)
project, emit [5]         exit 139
aggregate, emit [1, 0]    exit 139

The last variant is why this is more than a check on malformed input: that plan is valid Substrait. Arrow's vendored proto has no expression_references, so Arrow drops the grouping keys and derives an empty aggregate schema. The plan's valid [1, 0] mapping is then out of range for that empty schema. That grouping-key loss is #50634. A bounds check would turn the process crash into an error that reports the rejected [1, 0] mapping; fixing #50634 is still required for the plan to run.

Rechecked with PyArrow 26.0.0.dev244+g79e074ace, built from main at 79e074ace3a7c4ca26f211dfd84a1984ad53c362, on macOS arm64. The two positive out-of-range cases and the aggregate case still exit 139; the negative case now returns ArrowInvalid. The original PyArrow 25.0.1 reproduction crashed for the negative case as well on macOS arm64 and Linux x86-64. The source is cpp/src/arrow/engine/substrait/relation_internal.cc.

Reproducer — one file, pyarrow only
"""RelCommon.emit.output_mapping indices are used to index the input schema unchecked.

The control succeeds and the negative case is caught as `ArrowInvalid`. Run each variant in a process of its own because the other three terminate:
    for v in 0 1 2 3 4; do python3 emit_repro.py $v || echo "   exit $?"; done
"""
import json, sys
import pyarrow as pa
import pyarrow.substrait as ps
from pyarrow._substrait import _parse_json_plan

SCHEMA = pa.schema([pa.field("c0", pa.int64(), nullable=False),
                    pa.field("c1", pa.float64(), nullable=False)])

def provider(names, schema=None):
    return pa.table({"c0": [1], "c1": [2.5]}, schema=schema or SCHEMA)

def ref(i):
    return {"selection": {"directReference": {"structField": {"field": i} if i else {}},
                          "rootReference": {}}}

def emit(m):
    return {"common": {"emit": {"outputMapping": m}}}

def read(m=None):
    r = {"baseSchema": {"names": ["c0", "c1"], "struct": {
             "types": [{"i64": {"nullability": "NULLABILITY_REQUIRED"}},
                       {"fp64": {"nullability": "NULLABILITY_REQUIRED"}}],
             "nullability": "NULLABILITY_REQUIRED"}},
         "namedTable": {"names": ["t"]}}
    return {"read": dict(r, **(emit(m) if m else {}))}

def run(rel, names):
    plan = {"version": {"minorNumber": 102, "producer": "repro"},
            "relations": [{"root": {"input": rel, "names": names}}]}
    return ps.run_query(_parse_json_plan(json.dumps(plan).encode()), table_provider=provider)

# A plan that is valid Substrait: the grouping keys are where current Substrait puts them.
# Arrow's vendored proto has no expression_references, so it drops them (#50634) and the
# aggregate's output schema has no fields at all - which puts a correct mapping out of range.
AGG = {"aggregate": dict({"input": read(),
                          "groupings": [{"expressionReferences": [0, 1]}],
                          "groupingExpressions": [ref(0), ref(1)]}, **emit([1, 0]))}

VARIANTS = [
    ("read, emit [1, 0]",       lambda: run(read([1, 0]), ["c1", "c0"])),
    ("read, emit [5]",          lambda: run(read([5]), ["x"])),
    ("read, emit [-1]",         lambda: run(read([-1]), ["x"])),
    ("project, emit [5]",       lambda: run({"project": dict(
                                    {"input": read(), "expressions": [ref(0)]},
                                    **emit([5]))}, ["x"])),
    ("aggregate, emit [1, 0]",  lambda: run(AGG, ["c1", "c0"])),
]

label, case = VARIANTS[int(sys.argv[1])]
print("%-22s" % label, end=" ", flush=True)
try:
    print(case().read_all().schema.types)
except Exception as e:
    print(type(e).__name__ + ":", str(e).replace("\n", " ")[:70])
主要语言
C++
星标
17.2k
派生
4.3k
平均合并
4 天 3 小时
30 天内合并 PR
93

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/arrow 的其他 Issue

查看 apache/arrow 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。