Prompts of max_seq_len tokens fail to prefill: export_llm bounds the KV-cache token input at max_seq_len - 1 but publishes get_max_seq_len = max_seq_len
メンテナーはふだん 1 日以内に返信
評価
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 初心者へのやさしさ
- 20/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 停滞
調査の方向性
The bound comes from the dynamic-shape export in builder.py, reached through export_llm; compare it with the get_max_seq_len value the runner reads, and check text_llm_runner.cpp for how it chunks prefill. Done means a prompt of exactly get_max_seq_len tokens prefills in one call and longer prompts are chunked. Linked PR 23679 already targets this, so coordinate with it before starting.
索引モデルが issue の本文から書いたものです。
説明
🐛 Describe the bug
export_llm exports with a KV cache and dynamic shape (builder.py's default dynamic shapes) bound the token input of forward at max_seq_len - 1, but publish get_max_seq_len = max_seq_len. The text runner prefills in chunks of get_max_seq_len tokens, so a chunk of max_seq_len tokens does not fit and the prompt fails with a resize error.
With S = max_seq_length and C = max_context_length:
- S < C: on a fresh runner,
generate()fails for prompts of S to C - 1 tokens. - S == C (the default):
generate()refuses such prompts first ("Max seq length exceeded"), butprefill()/prefillPrompt()of exactly S tokens fails the same way.
The bound is S - 1 in every release from 1.1.0 to 1.5.1 and on main. 1.0.1 exports have bound S and work, also on the 1.5.1 runtime.
Expected: a prompt longer than get_max_seq_len is prefilled in chunks (the runner has chunked since #9805), and a prompt of exactly get_max_seq_len tokens is prefilled in one call.
Actual: on a fresh runner, the tested prompts of S - 1 tokens or fewer work; prompts of S to C - 1 tokens fail in Python (TextLLMRunner), C++ (llama_main) and Android (LlmModule).
To reproduce
pip install executorch==1.5.1 (torch 2.14.0, torchao 0.18.0). Model: stories110M.pt (sha256 cdef30b5108d1b5d1c232fcf7e2ea433cc6cdc473aa876af9d8d6ce83bf4b5e0) and the llama2.c tokenizer.model (sha256 9e556afd44213b6bd1be2b850ebbbd98f5481437a8021afaf58ee7fb1818d347).
echo '{"dim": 768, "multiple_of": 32, "n_heads": 12, "n_layers": 12, "norm_eps": 1e-05, "vocab_size": 32000}' > params.json
python -m executorch.extension.llm.export.export_llm base.model_class=stories110m \
base.checkpoint=stories110M.pt base.params=params.json \
model.use_kv_cache=True model.use_sdpa_with_kv_cache=True model.enable_dynamic_shape=True \
export.max_seq_length=128 export.max_context_length=512 backend.xnnpack.enabled=True \
export.output_name=stories110m_s128_c512.pte
# repro.py stories110m_s128_c512.pte tokenizer.model
import sys
from executorch.extension.llm.custom_ops import custom_ops # noqa: F401
from executorch.extension.llm.runner import GenerationConfig, TextLLMRunner
from executorch.runtime import Runtime
pte, tokenizer = sys.argv[1], sys.argv[2]
program = Runtime.get().load_program(pte)
print("get_max_seq_len", program.load_method("get_max_seq_len").execute([])[0])
print("forward token bound", program.metadata("forward").input_tensor_meta(0).sizes()[1])
config = GenerationConfig(echo=False, max_new_tokens=4, temperature=0.0)
TextLLMRunner(pte, tokenizer).generate(" ".join(["apple"] * 127), config, print) # 127 tokens: ok
TextLLMRunner(pte, tokenizer).generate(" ".join(["apple"] * 128), config, print) # 128 tokens: fails
Output (" apple" is one token in this tokenizer, so the prompts are exactly 127 and 128 tokens). stdout excerpt (the Python prints; the runner's own text and PyTorchObserver line are omitted):
get_max_seq_len 128
forward token bound 127
.
.
and
the
stderr, end:
[text_llm_runner.cpp:92] RSS after loading model: 0.000000 MiB (0 if unsupported)
[tensor_impl.cpp:161] Attempted to resize a bounded tensor with a maximum capacity of 127 elements to 128 elements.
[method.cpp:1246] Error resizing tensor at input 0
Traceback (most recent call last):
File "repro.py", line 15, in <module>
TextLLMRunner(pte, tokenizer).generate(" ".join(["apple"] * 128), config, print) # 128 tokens: fails
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
RuntimeError: Generation failed
The 127-token call generates; the 128-token call fails before the program runs.
Where it reproduces
The release, platform and cross-runtime checks below are stock export_llm exports (KV cache, dynamic shape), checked by one harness on GitHub Actions (prefill_check.py, run.sh, run). Where the Python runner is available (1.2.0 onward), it runs generate() with prompts of exactly S - 1, S, S + 1 and, where it fits, 2S + 1 tokens, counted with the runner's own C++ tokenizer, and prefill() of exactly S tokens; older releases are checked through forward(). The failing boundary calls report the resize error above.
| Release (torch) | Bound at S128/C512 | Result (stories110m, qwen3_0_6b, lfm2_350m; XNNPACK fp32; ubuntu-24.04) |
|---|---|---|
| 1.0.1 (2.9.1) | 128 | forward() accepts 128 tokens (no Python TextLLMRunner in this release) |
| 1.1.0 (2.10.0) | 127 | forward() fails at 128 tokens (no Python TextLLMRunner in this release) |
| 1.2.0, 1.3.1, 1.4.1, 1.5.1 (2.11.0 to 2.14.1), nightly 1.6.0.dev20261008 | 127 | generate() ok at 127, fails at 128, 129, 257; prefill() of 128 fails; 257 tokens as 127 + 127 + 3 prefill() pieces ok |
| Also checked | Result |
|---|---|
ubuntu-24.04-arm, macos-15 runners; 8da4w; MLX (backend.mlx.enabled=True) |
same failure |
| S2048/C8192 | bound 2047; generate() fails at 2048, 2049, 4097 |
| S512/C512 | bound 511; prefill() of 512 fails (generate() refuses 512 by the context check) |
Vulkan (backend.vulkan.enabled=True, fp32 and 8da4w with force_fp16; main built from source, run) |
bound 127; same failure on SwiftShader, and with llama_main on a Mali GPU (MediaTek MT6991) and an Adreno GPU (Snapdragon 8 Elite Gen 5, SM8850) |
Android 16: LlmModule from executorch-android 1.5.1 on the SM8850 and MT6991; llama_main (XNNPACK) built from main on both phones and an arm64 emulator (run) |
127 tokens ok; 128, 129, 257 fail (ExecuTorch Error 0x10 from LlmModule, resize error in logcat) |
| qwen3_0_6b and lfm2_350m on the phones (XNNPACK fp32 and 8da4w, Vulkan 8da4w; phone_models.sh, device.sh, run outside CI) | same failure, except lfm2_350m on Vulkan, which aborts on a separate delegate error (Vulkan graph declares 2 inputs and 11 outputs, but the delegate call supplied 3 arguments) at any length that reaches the delegate |
| Exported with 1.0.1, run on 1.5.1 and the nightly | bound 128; the tested generate() and prefill() cases pass on both, so an S-token input works when the file's bound is S |
Not affected:
- The MLX backend's own HF exporter (
backends/mlx/examples/llm/export_llm_hf.py): its SmolLM2-135M export has bound 128 and works at 128, 129 and 257 tokens (checked locally). - The Qualcomm llama pipeline:
export_llmrejects dynamic shape with QNN (export_llama_lib.py#L1119-L1125) andqnn_llama_runnertakes its chunk size from the program's static shapes. On the SM8850 it prefills 32 to 258 tokens in 1 to 9 chunks of 32. - MediaTek's llama runner, which prefills in batches of its static prompt model (mtk_llama_runner.cpp#L151-L170). With a NeuroPilot export of Qwen2.5-0.5B-Instruct for the Dimensity 9400 (run), it exits 0 and prints response text for prompts of 126 to 257 tokens on the MT6991. (That export drops the shared-weights compile spec, because the public NeuroPilot SDK that CI installs has no
extract_shared_data; its output after more than one 128-token batch was not coherent, which I did not investigate.)
Every cell (21 release rows, 14 platform rows, 4 cross rows)
Releases (ubuntu-24.04, S128/C512, fp32)
| Version | Model | Bound | forward() tokens |
generate() prompt tokens |
prefill() of S |
2S + 1 in pieces |
|---|---|---|---|---|---|---|
| 1.0.1+cpu | lfm2_350m | 128 | 127 ok, 128 ok, 129 fails | no TextLLMRunner binding |
||
| 1.0.1+cpu | qwen3_0_6b | 128 | 127 ok, 128 ok, 129 fails | no TextLLMRunner binding |
||
| 1.0.1+cpu | stories110m | 128 | 127 ok, 128 ok, 129 fails | no TextLLMRunner binding |
||
| 1.1.0+cpu | lfm2_350m | 127 | 127 ok, 128 fails, 129 fails | no TextLLMRunner binding |
||
| 1.1.0+cpu | qwen3_0_6b | 127 | 127 ok, 128 fails, 129 fails | no TextLLMRunner binding |
||
| 1.1.0+cpu | stories110m | 127 | 127 ok, 128 fails, 129 fails | no TextLLMRunner binding |
||
| 1.2.0+cpu | lfm2_350m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.2.0+cpu | qwen3_0_6b | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.2.0+cpu | stories110m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.3.1+cpu | lfm2_350m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.3.1+cpu | qwen3_0_6b | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.3.1+cpu | stories110m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.4.1+cpu | lfm2_350m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.4.1+cpu | qwen3_0_6b | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.4.1+cpu | stories110m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.5.1+cpu | lfm2_350m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.5.1+cpu | qwen3_0_6b | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.5.1+cpu | stories110m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.6.0.dev20261008+cpu | lfm2_350m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.6.0.dev20261008+cpu | qwen3_0_6b | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| 1.6.0.dev20261008+cpu | stories110m | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
Platforms, sizes, quantization, MLX
| Runner | Version | Model | Shape | Quant | Backend | Bound | forward() tokens |
generate() prompt tokens |
prefill() of S |
2S + 1 in pieces |
|---|---|---|---|---|---|---|---|---|---|---|
| Darwin arm64 | 1.5.1 | qwen3_0_6b | S128/C512 | fp32 | MLX | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Darwin arm64 | 1.5.1 | stories110m | S128/C512 | fp32 | MLX | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Darwin arm64 | 1.6.0.dev20261008 | qwen3_0_6b | S128/C512 | fp32 | MLX | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Darwin arm64 | 1.6.0.dev20261008 | stories110m | S128/C512 | fp32 | MLX | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Darwin arm64 | 1.5.1 | lfm2_350m | S128/C512 | fp32 | XNNPACK | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Darwin arm64 | 1.5.1 | qwen3_0_6b | S128/C512 | fp32 | XNNPACK | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Darwin arm64 | 1.6.0.dev20261008 | qwen3_0_6b | S128/C512 | fp32 | XNNPACK | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Linux x86_64 | 1.0.1+cpu | qwen3_0_6b | S512/C512 | fp32 | XNNPACK | 512 | 511 ok, 512 ok, 513 fails | no TextLLMRunner binding |
||
| Linux x86_64 | 1.5.1+cpu | qwen3_0_6b | S128/C512 | 8da4w | XNNPACK | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Linux x86_64 | 1.5.1+cpu | qwen3_0_6b | S2048/C8192 | 8da4w | XNNPACK | 2047 | 2047 ok, 2048 fails, 2049 fails | 2047 ok, 2048 fails, 2049 fails, 4097 fails | fails | 2047+2047+3 ok |
| Linux x86_64 | 1.5.1+cpu | qwen3_0_6b | S512/C512 | fp32 | XNNPACK | 511 | 511 ok, 512 fails, 513 fails | 511 ok, 512 refused (context), 513 refused (context) | fails | |
| Linux aarch64 | 1.5.1+cpu | qwen3_0_6b | S128/C512 | fp32 | XNNPACK | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Linux aarch64 | 1.6.0.dev20261008+cpu | qwen3_0_6b | S128/C512 | fp32 | XNNPACK | 127 | 127 ok, 128 fails, 129 fails | 127 ok, 128 fails, 129 fails, 257 fails | fails | 127+127+3 ok |
| Linux x86_64 | 1.6.0.dev20261008+cpu | stories110m | S2048/C8192 | 8da4w | XNNPACK | 2047 | 2047 ok, 2048 fails, 2049 fails | 2047 ok, 2048 fails, 2049 fails, 4097 fails | fails | 2047+2047+3 ok |
Exported with 1.0.1, run on a newer runtime
| Runtime | Model | Bound | forward() tokens |
generate() prompt tokens |
prefill() of S |
2S + 1 in pieces |
|---|---|---|---|---|---|---|
| 1.5.1+cpu | qwen3_0_6b | 128 | 127 ok, 128 ok, 129 fails | 127 ok, 128 ok, 129 ok, 257 ok | ok | 128+128+1 ok, same text |
| 1.5.1+cpu | stories110m | 128 | 127 ok, 128 ok, 129 fails | 127 ok, 128 ok, 129 ok, 257 ok | ok | 128+128+1 ok, same text |
| 1.6.0.dev20261008+cpu | qwen3_0_6b | 128 | 127 ok, 128 ok, 129 fails | 127 ok, 128 ok, 129 ok, 257 ok | ok | 128+128+1 ok, same text |
| 1.6.0.dev20261008+cpu | stories110m | 128 | 127 ok, 128 ok, 129 fails | 127 ok, 128 ok, 129 ok, 257 ok | ok | 128+128+1 ok, same text |
Root cause
Links at main 9875560; builder.py and the runner files below are unchanged on main as of f2575be.
- builder.py#L148-L151 bounds the token dimension of KV-cache exports at
max_seq_len - 1, "due to export limitation (same as non-kv-cache case above)", while builder.py#L119 and export_llama_lib.py#L1939 publishget_max_seq_len = max_seq_len. - llm_runner_helper.cpp#L320-L324 gives
TextPrefillerthat value as its chunk size, and text_prefiller.cpp#L52-L60 prefills in chunks of it. - text_llm_runner.cpp#L142-L170 lets prompts longer than
max_seq_lenthrough to the prefiller when S < C (1.1.0 and 1.2.0 check only the context left).TextLLMRunner::prefillhas no such check, soprefill()of S tokens reaches the program in every case.
History of the KV-cache bound:
| Change | Merged | KV-cache token bound |
|---|---|---|
| #9805 chunked prefill in the runner | 2025-04-01 | chunk = get_max_seq_len |
| #14737 | 2025-10-02 | max_seq_len - 1, because update_cache checked start_pos + seq_len < cache size |
| #15084 "Fix max seq length bug" | 2025-10-15 | max_seq_len; the check became <=, and test_builder.py expected max_seq_len |
| #15951 "Decompose after export in export_llama" | 2025-12-09 | max_seq_len - 1 again, "same as non-kv-cache case above", and the test changed back |
#15951 is about unwrap_tensor_subclass and LoRA. The gist its comment cites shows the guard only for the graph without a KV cache (L['tokens'].size()[1] != 128). With the KV-cache bound set back to max_seq_len on main, stories110m, qwen3_0_6b and lfm2_350m (S128/C512), qwen3_0_6b 8da4w (S2048/C8192) and S512/C512 exports all export with bound S and pass the checks above on ubuntu-24.04 and macos-15; the no-KV-cache branch still needs its - 1.
Proposed fix
A PR follows with two parts:
builder.py: the KV-cache branch bounds the token input atmax_seq_lenagain (the no-KV-cache branch keeps- 1), andtest_builder.pyexpects it, as after #15084.- Runner: files already exported by 1.1 to 1.5 keep the
- 1bound, socreate_text_llm_runnercapsTextPrefiller's chunk at the upper bound of the method's token input (frommethod_meta) for KV-cache, dynamic-shape models.
Workaround for existing files: send the prompt in pieces of at most the bound (prefill() / prefillPrompt(), then generate()).
Related: software-mansion/react-native-executorch#1491 clamps the chunk to method_meta.input_tensor_meta(0) for the same reason; optimum-executorch bounds its export at min(max_seq_len, sliding_window) - 1 (huggingface/optimum-executorch#218).
Versions
collect_env (macOS repro above)
Collecting environment information...
PyTorch version: 2.14.0
Is debug build: False
CUDA used to build PyTorch: None
ROCM used to build PyTorch: N/A
OS: macOS 26.5.2 (arm64)
GCC version: Could not collect
Clang version: 21.0.0 (clang-2100.1.1.101)
CMake version: version 4.4.0
Libc version: N/A
Python version: 3.12.13 (main, Jun 23 2026, 15:44:24) [Clang 22.1.3 ] (64-bit runtime)
Python platform: macOS-26.5.2-arm64-arm-64bit
Is CUDA available: False
CUDA runtime version: No CUDA
CUDA_MODULE_LOADING set to: N/A
GPU models and configuration: No CUDA
Nvidia driver version: No CUDA
cuDNN version: No CUDA
Is XPU available: False
HIP runtime version: N/A
MIOpen runtime version: N/A
Is XNNPACK available: False
Caching allocator config: N/A
CPU:
Apple M5
Versions of relevant libraries:
[pip3] Could not collect
[conda] Could not collect
executorch 1.5.1
numpy 2.5.3
pytorch-tokenizers 1.5.0
torch 2.14.0
torchao 0.18.0
Also on GitHub Actions: executorch 1.0.1 to 1.5.1 and nightly 1.6.0.dev20261008, each with the torch it pins, on ubuntu-24.04, ubuntu-24.04-arm and macos-15. executorch-android 1.5.1 on Android 16.
cc @mergennachin @iseeyuan @lucylq @helunwencser @tarun292 @kimishpatel @jackzhxng
- 主要言語
- Python
- スター
- 5.1k
- フォーク
- 1.2k
- 平均マージ
- 2日 9時間
- マージ済み PR(30日)
- 661
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
pytorch/executorch のほかの issue
-
enhancement triaged
難易度 2/5 半日 初心者へのやさしさ 68/100
pytorch/executorch#21640 ·
メンテナーはふだん 1 日以内に返信
-
[v1.6.0] Release Schedule and Tracker対応中かも @JacobSzwejbka が 2 日前に担当しました。 オープン
pytorch/executorch#23608 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
QNN sharded LLM export fails: ResolveDebugHandle is not the last edge pass (dep_table entry overwritten by SplitGraph registration)対応中かも @psiddh が 3 日前に担当しました。 オープンbug module: llm module: qnn triaged
pytorch/executorch#23580 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
pytorch/executorch#23553 ·
メンテナーはふだん 1 日以内に返信
-
module: ci module: samsung unstable
難易度 2/5 1〜3時間 初心者へのやさしさ 50/100
pytorch/executorch#23507 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
pytorch/executorch の issue をすべて見る
似ている issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 83/100
PedestrianDynamics/pyFDS-Evac#766 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1〜3時間 初心者へのやさしさ 91/100
alchaincyf/nuwa-skill#86 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 2 日以内に返信
-
Docs Needs Triage
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
pandas-dev/pandas#71055 ·
メンテナーはふだん 1 日以内に返信
-
[Bug]: graphify reads files that git's global ignore file hides対応中かも @smngvlkz が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
Graphify-Labs/graphify#4335 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信