Runtime accepts a channels-last input for a contiguous method and reads it as contiguous (silent wrong output)
维护者通常 1 天内回复
@JakeStevens 已经在做这个了。
开始于 2026年8月14日。
评估
这个 Issue 还没有评估数据。
描述
The runtime validates input layout — a transposed 2-D tensor is rejected with
Input 0 for method forward should be contiguous or channels-last. But a
channels-last 4-D tensor is accepted even when the method was exported with a
contiguous input, and its data is then read in contiguous order. No error, no
warning, wrong numbers.
Repro (executorch 1.4.0, torch 2.13, macOS arm64, portable kernels — the XNNPACK
partitioner makes no difference):
import torch
from executorch.exir import to_edge_transform_and_lower
from executorch.runtime import Runtime
class M(torch.nn.Module):
def forward(self, x):
return x * 2.0
contig = torch.arange(12, dtype=torch.float32).reshape(1, 3, 2, 2)
chlast = contig.to(memory_format=torch.channels_last)
assert torch.equal(contig, chlast)
ep = torch.export.export(M(), (contig,))
open("/tmp/r.pte", "wb").write(
to_edge_transform_and_lower(ep, partitioner=[]).to_executorch().buffer)
m = Runtime.get().load_program("/tmp/r.pte").load_method("forward")
print("expected :", (contig * 2).flatten().tolist())
print("contiguous :", m.execute([contig])[0].flatten().tolist())
print("channels_last:", m.execute([chlast])[0].flatten().tolist())
expected : [0.0, 2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0]
contiguous : [0.0, 2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0]
channels_last: [0.0, 8.0, 16.0, 2.0, 10.0, 18.0, 4.0, 12.0, 20.0, 6.0, 14.0, 22.0]
Two tensors that torch.equal reports as equal produce different results.
Why this is worth fixing rather than documenting. Channels-last is what you
get from the most common way to build an image input:
torch.from_numpy(hwc_array.transpose(2, 0, 1)[None]) # NCHW shape, NHWC strides
np.transpose returns a view and astype keeps memory order by default
(order='K'), so this is easy to hit without ever naming channels_last.
The failure mode is what makes it expensive. torch.randn inputs are contiguous
and match the eager model bit-exactly, so a parity check passes; only real images
diverge. I spent an afternoon bisecting eager -> ep.module() ->
edge.exported_program().module() -> .pte on a Depth-Anything-V2 export,
concluding the conversion was broken, before finding the input was the difference.
The first three stages honour strides and agree; only the runtime disagrees.
Either honouring the strides or rejecting the mismatch would have cost me nothing
to debug. Rejecting seems in keeping with the existing check — the message just
needs to compare the input's layout against the method's expected layout rather
than accept channels-last unconditionally.
cc @mergennachin @kimishpatel @iseeyuan @larryliu0820 @JacobSzwejbka @lucylq
- 主要语言
- Python
- 星标
- 5k
- 派生
- 1.2k
- 平均合并
- 2 天 5 小时
- 30 天内合并 PR
- 491
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
pytorch/executorch 的其他 Issue
-
enhancement triaged
难度 2/5 半天 新手友好度 68/100
pytorch/executorch#21640 ·
维护者通常 1 天内回复
-
enhancement module: examples
难度 5/5 一周以上 新手友好度 20/100
pytorch/executorch#23164 · 7 条评论 · 1 个 reaction ·
维护者通常 1 天内回复
-
Qualcomm: 8-bit per-channel weight scales are floored at the 16-bit eps, and the HTP miscomputes near-zero channels可能已有人在做 @psiddh 于 3 天前认领。 未关闭module: qnn partner: qualcomm
pytorch/executorch#23160 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[cpu kernels] native_layer_norm: layer_norm_scalar returns NaN on large-mean rows; Half/BF16 at N>=256 slow after #23153可能已有人在做 @JakeStevens 于 3 天前认领。 未关闭module: kernels
pytorch/executorch#23159 · 2 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
module: vulkan
难度 3/5 1-2 天 新手友好度 66/100
pytorch/executorch#23158 ·
维护者通常 1 天内回复
查看 pytorch/executorch 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 86/100
维护者通常 1 天内回复
-
难度 2/5 1-2 天 新手友好度 70/100
-
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 7 天内回复
-
难度 2/5 1-3 小时 新手友好度 70/100
lmstudio-ai/mlx-engine#376 ·
-
难度 2/5 1-3 小时 新手友好度 72/100
pyiron/bagofholding#166 ·