Runtime accepts a channels-last input for a contiguous method and reads it as contiguous (silent wrong output)
メンテナーはふだん 1 日以内に返信
@JakeStevens がすでに取り組んでいます。
2026年8月14日 から。
評価
この issue はまだ評価されていません。
説明
The runtime validates input layout — a transposed 2-D tensor is rejected with
Input 0 for method forward should be contiguous or channels-last. But a
channels-last 4-D tensor is accepted even when the method was exported with a
contiguous input, and its data is then read in contiguous order. No error, no
warning, wrong numbers.
Repro (executorch 1.4.0, torch 2.13, macOS arm64, portable kernels — the XNNPACK
partitioner makes no difference):
import torch
from executorch.exir import to_edge_transform_and_lower
from executorch.runtime import Runtime
class M(torch.nn.Module):
def forward(self, x):
return x * 2.0
contig = torch.arange(12, dtype=torch.float32).reshape(1, 3, 2, 2)
chlast = contig.to(memory_format=torch.channels_last)
assert torch.equal(contig, chlast)
ep = torch.export.export(M(), (contig,))
open("/tmp/r.pte", "wb").write(
to_edge_transform_and_lower(ep, partitioner=[]).to_executorch().buffer)
m = Runtime.get().load_program("/tmp/r.pte").load_method("forward")
print("expected :", (contig * 2).flatten().tolist())
print("contiguous :", m.execute([contig])[0].flatten().tolist())
print("channels_last:", m.execute([chlast])[0].flatten().tolist())
expected : [0.0, 2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0]
contiguous : [0.0, 2.0, 4.0, 6.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0]
channels_last: [0.0, 8.0, 16.0, 2.0, 10.0, 18.0, 4.0, 12.0, 20.0, 6.0, 14.0, 22.0]
Two tensors that torch.equal reports as equal produce different results.
Why this is worth fixing rather than documenting. Channels-last is what you
get from the most common way to build an image input:
torch.from_numpy(hwc_array.transpose(2, 0, 1)[None]) # NCHW shape, NHWC strides
np.transpose returns a view and astype keeps memory order by default
(order='K'), so this is easy to hit without ever naming channels_last.
The failure mode is what makes it expensive. torch.randn inputs are contiguous
and match the eager model bit-exactly, so a parity check passes; only real images
diverge. I spent an afternoon bisecting eager -> ep.module() ->
edge.exported_program().module() -> .pte on a Depth-Anything-V2 export,
concluding the conversion was broken, before finding the input was the difference.
The first three stages honour strides and agree; only the runtime disagrees.
Either honouring the strides or rejecting the mismatch would have cost me nothing
to debug. Rejecting seems in keeping with the existing check — the message just
needs to compare the input's layout against the method's expected layout rather
than accept channels-last unconditionally.
cc @mergennachin @kimishpatel @iseeyuan @larryliu0820 @JacobSzwejbka @lucylq
- 主要言語
- Python
- スター
- 5k
- フォーク
- 1.2k
- 平均マージ
- 2日 5時間
- マージ済み PR(30日)
- 491
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
pytorch/executorch のほかの issue
-
enhancement triaged
難易度 2/5 半日 初心者へのやさしさ 68/100
pytorch/executorch#21640 ·
メンテナーはふだん 1 日以内に返信
-
enhancement module: examples
難易度 5/5 1週間以上 初心者へのやさしさ 20/100
pytorch/executorch#23164 · コメント 7 件 · リアクション 1 件 ·
メンテナーはふだん 1 日以内に返信
-
Qualcomm: 8-bit per-channel weight scales are floored at the 16-bit eps, and the HTP miscomputes near-zero channels対応中かも @psiddh が 2 日前に担当しました。 オープンmodule: qnn partner: qualcomm
pytorch/executorch#23160 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
[cpu kernels] native_layer_norm: layer_norm_scalar returns NaN on large-mean rows; Half/BF16 at N>=256 slow after #23153対応中かも @JakeStevens が 2 日前に担当しました。 オープンmodule: kernels
pytorch/executorch#23159 · コメント 2 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
module: vulkan
難易度 3/5 1〜2日 初心者へのやさしさ 66/100
pytorch/executorch#23158 ·
メンテナーはふだん 1 日以内に返信
pytorch/executorch の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
-
bug
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
qgis/QGIS-Plugins-Website#459 ·
-
bug severity:medium
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 2 日以内に返信
-
bot-found bug priority: P3
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
madenvel/KalinkaPlayer#179 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
ls1intum/edutelligence#1098 ·
メンテナーはふだん 1 日以内に返信