Metal backend: generated kernels index a packed copy of a non-packed view out of bounds (wrong results, crash in MPSGraph)
@abdelaziz-mahdy ya está trabajando en esto.
Desde el 22/9/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
🐛 Describe the bug
With inductor's default layout optimization, models built on chunk + cat along C (the YOLO C2f block) return wrong values on the Metal backend, and yolov8n crashes inside MPSGraph.
aoti_torch__reinterpret_tensor replaces a view that is not densely packed with a packed copy (materialize_packed). Kernels generated by inductor never look at the tensor they are handed: they index with the strides they were compiled with, and they write through such views as well as read them. In channels-last a chunk along C is non-packed, and a cat is filled by writing through views of it:
buf52 = reinterpret_tensor_wrapper(buf55, 4, {1, 32, 80, 80}, {819200, 1, 10240, 128}, 32) // alias
buf55 has 819200 elements and the copy that stands in for buf52 has 204800. The kernel writing "into buf52" indexes up to 819200 into that copy, which is an out-of-bounds GPU access, and nothing reaches buf55. #22957 lists the lost writes under "Not fixed here" (linear_chunk_cat_last_dim is skipped for it); the out-of-bounds part is new, and it is what takes yolov8n down: 10 out of 10 runs crash with EXC_BAD_ACCESS in -[MPSGraphExecutable runInternalWithDevice:...] -> getFuncOp -> mlir::SymbolTable::lookupSymbolIn, called from aoti_torch_mps_convolution, on a pointer that reads as two floats. With torch._inductor.config.layout_optimization = False there are no such views and the same model runs correctly.
Repro. The C2f block reduced to 1x1 convs, so that it only needs matmuls. For MODULE_REGISTRY in backends/apple/metal/tests/test_modules.py:
class PointwiseC2f(nn.Module):
class Inner(nn.Module):
def __init__(self, channels: int):
super().__init__()
self.conv1 = nn.Conv2d(channels, channels, kernel_size=1)
self.conv2 = nn.Conv2d(channels, channels, kernel_size=1)
def forward(self, x):
return x + self.conv2(torch.relu(self.conv1(x)))
def __init__(self):
super().__init__()
self.conv_in = nn.Conv2d(16, 16, kernel_size=1)
self.inner = nn.ModuleList(PointwiseC2f.Inner(8) for _ in range(2))
self.conv_out = nn.Conv2d(32, 16, kernel_size=1)
def forward(self, x):
parts = list(self.conv_in(x).chunk(2, dim=1))
for block in self.inner:
parts.append(block(parts[-1]))
return self.conv_out(torch.cat(parts, dim=1))
MODULE_REGISTRY["pointwise_c2f"] = {
"model_class": PointwiseC2f,
"input_shapes": [(2, 16, 8, 8)],
"description": "C2f block whose cat is filled through non-packed channels-last views",
}
Output mismatch - max_atol=0.586, max_rtol=1.64 in float32, 0.755 / 1.71 in bfloat16.
Versions
ExecuTorch: #22957 @ 89dbe34077 (main @ 9b91b43098 plus that PR)
PyTorch version: 2.14.0
OS: macOS 27.0 (26A428), arm64, Apple M2 Pro
Clang version: 21.0.0 (clang-2100.3.34.2)
CMake version: 4.4.3
Python version: 3.10.11
- Lenguaje dominante
- Python
- Estrellas
- 5k
- Forks
- 1.2k
- Merge medio
- 2 d 12 h
- PR fusionados (30 d)
- 588
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de pytorch/executorch
-
enhancement triaged
Dificultad 2/5 Medio día Aptitud para principiantes 68/100
pytorch/executorch#21640 ·
-
enhancement module: cuda module: examples module: mlx module: vulkan module: xnnpack
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
pytorch/executorch#23131 · 1 reacción ·
-
bug module: qnn module: quantization
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
pytorch/executorch#23108 ·
-
good first issue module: examples module: webgpu
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
pytorch/executorch#23099 ·
-
good first issue module: examples module: vulkan
Dificultad 5/5 Más de una semana Aptitud para principiantes 30/100
pytorch/executorch#23098 ·
Todos los issues de pytorch/executorch
Issues similares
-
agent-ready documentation needs-triage
Dificultad 1/5 1-3 horas Aptitud para principiantes 88/100
-
documentation
Dificultad 1/5 Menos de una hora Aptitud para principiantes 91/100
-
workflow-status page template still says reusable workflows are "triggered only by workflow_call:" Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
-
Add https://search.jeremyh.xyz/ Abiertoinstance instance add
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
searxng/searx-instances#939 · 1 comentario ·
-
area-deployment area-integrations triage:bot-seen
Dificultad 2/5 Medio día Aptitud para principiantes 86/100