Metal backend: generated kernels index a packed copy of a non-packed view out of bounds (wrong results, crash in MPSGraph)
@abdelaziz-mahdy 已经在做这个了。
开始于 2026年9月22日。
评估
这个 Issue 还没有评估数据。
描述
🐛 Describe the bug
With inductor's default layout optimization, models built on chunk + cat along C (the YOLO C2f block) return wrong values on the Metal backend, and yolov8n crashes inside MPSGraph.
aoti_torch__reinterpret_tensor replaces a view that is not densely packed with a packed copy (materialize_packed). Kernels generated by inductor never look at the tensor they are handed: they index with the strides they were compiled with, and they write through such views as well as read them. In channels-last a chunk along C is non-packed, and a cat is filled by writing through views of it:
buf52 = reinterpret_tensor_wrapper(buf55, 4, {1, 32, 80, 80}, {819200, 1, 10240, 128}, 32) // alias
buf55 has 819200 elements and the copy that stands in for buf52 has 204800. The kernel writing "into buf52" indexes up to 819200 into that copy, which is an out-of-bounds GPU access, and nothing reaches buf55. #22957 lists the lost writes under "Not fixed here" (linear_chunk_cat_last_dim is skipped for it); the out-of-bounds part is new, and it is what takes yolov8n down: 10 out of 10 runs crash with EXC_BAD_ACCESS in -[MPSGraphExecutable runInternalWithDevice:...] -> getFuncOp -> mlir::SymbolTable::lookupSymbolIn, called from aoti_torch_mps_convolution, on a pointer that reads as two floats. With torch._inductor.config.layout_optimization = False there are no such views and the same model runs correctly.
Repro. The C2f block reduced to 1x1 convs, so that it only needs matmuls. For MODULE_REGISTRY in backends/apple/metal/tests/test_modules.py:
class PointwiseC2f(nn.Module):
class Inner(nn.Module):
def __init__(self, channels: int):
super().__init__()
self.conv1 = nn.Conv2d(channels, channels, kernel_size=1)
self.conv2 = nn.Conv2d(channels, channels, kernel_size=1)
def forward(self, x):
return x + self.conv2(torch.relu(self.conv1(x)))
def __init__(self):
super().__init__()
self.conv_in = nn.Conv2d(16, 16, kernel_size=1)
self.inner = nn.ModuleList(PointwiseC2f.Inner(8) for _ in range(2))
self.conv_out = nn.Conv2d(32, 16, kernel_size=1)
def forward(self, x):
parts = list(self.conv_in(x).chunk(2, dim=1))
for block in self.inner:
parts.append(block(parts[-1]))
return self.conv_out(torch.cat(parts, dim=1))
MODULE_REGISTRY["pointwise_c2f"] = {
"model_class": PointwiseC2f,
"input_shapes": [(2, 16, 8, 8)],
"description": "C2f block whose cat is filled through non-packed channels-last views",
}
Output mismatch - max_atol=0.586, max_rtol=1.64 in float32, 0.755 / 1.71 in bfloat16.
Versions
ExecuTorch: #22957 @ 89dbe34077 (main @ 9b91b43098 plus that PR)
PyTorch version: 2.14.0
OS: macOS 27.0 (26A428), arm64, Apple M2 Pro
Clang version: 21.0.0 (clang-2100.3.34.2)
CMake version: 4.4.3
Python version: 3.10.11
- 主要语言
- Python
- 星标
- 5k
- 派生
- 1.2k
- 平均合并
- 2 天 12 小时
- 30 天内合并 PR
- 588
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
pytorch/executorch 的其他 Issue
-
enhancement triaged
难度 2/5 半天 新手友好度 68/100
pytorch/executorch#21640 ·
-
enhancement module: examples
pytorch/executorch#23164 · 7 条评论 · 1 个 reaction ·
-
module: qnn partner: qualcomm
pytorch/executorch#23160 · 1 条评论 · 已指派 1 人 ·
-
module: kernels
pytorch/executorch#23159 · 2 条评论 · 已指派 1 人 ·
-
module: vulkan
pytorch/executorch#23158 ·
查看 pytorch/executorch 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 75/100
-
hcocena 未关闭policies-accepted pre-review precheck-passed
难度 1/5 1 小时以内 新手友好度 88/100
Bioconductor/BiocContributions#214 · 5 条评论 ·
-
难度 1/5 1 小时以内 新手友好度 92/100
TencentCloud/Octop#1169 · 1 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 70/100
521xueweihan/HelloGitHub#3778 ·
-
The version checker's trailing attribute region has no control for a less-than inside a quoted value 未关闭area: dashboard area: tests bug perceived difficulty: 2 python
难度 2/5 1-3 小时 新手友好度 84/100
Nitjsefnie-Harness-Commons/daedalus#1105 · 1 条评论 ·